Courseiva

CCNA Implementing Cicd Questions

40 questions · Implementing Cicd topic · All types, answers revealed

1
MCQmedium

Which of the following is a fundamental principle of implementing effective CI/CD for Databricks workflows?

A.Deploying code from a local laptop directly to the production cluster.
B.Using environment variables to manage service endpoints and authentication secrets.
C.Merging all development code into the main branch every few months.
D.Ignoring unit testing until the code is fully deployed in production.
AnswerB

Environment variables decouple code from the underlying infrastructure, allowing developers to manage configuration dynamically. This is a best practice in CI/CD as it enables the same codebase to run in any environment by simply injecting the appropriate parameters, reducing the risk of hardcoding secrets or environment-specific values.

Why this answer

Effective CI/CD requires separating code logic from environment-specific configurations. By parameterizing job settings and using environment-specific variables, teams ensure that the same code can be deployed safely across development, testing, and production. This practice minimizes errors during promotion and allows for automated testing cycles, which are essential for maintaining a robust, stable, and scalable data platform in any enterprise Databricks environment.

Exam trap

Candidates frequently select answers related to manual UI configuration or hardcoding values in code, ignoring that CI/CD best practices require environment abstraction through parameterization and secret management.

2
MCQhard

A CI/CD pipeline uses the Databricks CLI to deploy a job to production. The pipeline must ensure that the job configuration is identical across environments except for the cluster size, which differs between staging and production. Which approach best supports this requirement?

A.Define the job in a JSON file with placeholders, and use environment variables in the CI/CD pipeline to substitute cluster size values before deployment.
B.Maintain separate JSON files for staging and production, and manually update both when the job logic changes.
C.Store the job definition in a Databricks notebook and use widgets to parameterize the cluster size at runtime.
D.Use the Databricks REST API to create the job in production, then manually edit the cluster size in the Databricks UI after deployment.
AnswerA

Using a templated JSON file with environment-specific variables allows the pipeline to inject the correct cluster size at deployment time. This keeps the job definition consistent across environments while allowing controlled differences. It is a common pattern for infrastructure-as-code and aligns with CI/CD best practices for Databricks jobs.

Why this answer

Templating the job JSON with environment variables allows the pipeline to substitute environment-specific values like cluster size while keeping the rest of the configuration identical. This ensures consistency and reduces drift, as the same template is used for all environments, and differences are explicitly managed through variables.

Exam trap

The trap here is confusing runtime parameters (like widgets) with deployment-time configuration, which must be injected before the job is created or updated.

3
MCQmedium

A data engineering team uses Databricks Asset Bundles (DABs) to manage a job that must deploy to both a staging and a production workspace. The team wants to avoid hardcoding workspace-specific values such as the cluster ID and the storage path in databricks.yml. Which approach should they use?

A.Run databricks bundle deploy once, then use the Databricks REST API to patch the job's cluster ID and storage path after each deployment.
B.Define variables in the bundle and set target-specific values in databricks.yml under targets, then deploy with databricks bundle deploy -t <target>.
C.Store the cluster ID and storage path in a Databricks secret scope and reference them with dbutils.secrets.get inside the job's notebook code.
D.Create a separate Git branch for each workspace and manually edit the cluster ID and storage path in databricks.yml before each deployment.
AnswerB

DABs support variables and targets in databricks.yml. You declare variables once and override their values per target (for example, staging and prod), so the same bundle definition deploys correctly to each workspace. Running databricks bundle deploy -t <target> selects the target and applies its variable values, eliminating hardcoded cluster IDs and paths while keeping one source of truth.

Why this answer

Databricks Asset Bundles use a declarative databricks.yml with variables and targets. Declaring variables once and overriding them per target lets one bundle deploy consistently to staging and production without hardcoded workspace-specific values. Selecting the target at deploy time applies the correct values, keeps a single source of truth in Git, and avoids drift from manual edits or post-deployment patching.

Exam trap

The trap here is assuming environment-specific values must be embedded in the bundle file or patched afterward, when DABs targets and variables are designed precisely to parameterize them at deploy time.

4
Multi-Selecthard

A platform team is setting up a CI/CD pipeline that deploys Databricks jobs and notebooks from a Git repository. They must ensure deployments are secure and auditable. (Choose two.)

Select 2 answers
A.Grant the deployment identity only the permissions required to manage the specific jobs and notebooks it deploys.
B.Authenticate the pipeline with a service principal and store its OAuth secret in the CI/CD system's secret store.
C.Disable workspace audit logs during deployments to reduce log volume and improve pipeline speed.
D.Use a different engineer's personal account for each environment so responsibility is clear.
E.Commit the workspace personal access token to the repository so all pipeline runs use the same credential.
AnswersA, B

Least-privilege permissions limit the blast radius if the pipeline credential is compromised and keep audit logs focused on the resources the pipeline should touch. This directly supports the security requirement while preserving traceability of what the deployment identity changed.

Why this answer

Secure, auditable deployments rely on a dedicated non-human identity with narrowly scoped permissions. A service principal whose secret lives in the CI/CD secret store provides traceable, rotatable credentials, and least-privilege grants limit what a compromised pipeline could do while keeping audit records meaningful.

Exam trap

The trap here is treating any working credential, such as a committed personal access token, as acceptable for automation when the requirement is specifically secure and auditable deployment.

5
Multi-Selectmedium

A team is setting up a CI/CD pipeline that deploys Databricks Asset Bundles to production. They want the pipeline to be secure and to fail fast before any production resources are changed. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Run the deployment on every push to any branch so problems are discovered as early as possible in feature development.
B.Authenticate the deployment using a service principal with credentials injected from the CI system's secret manager at runtime.
C.Run databricks bundle validate in an earlier pipeline stage so configuration errors are caught before the deploy step executes.
D.Store the service principal secret as a plaintext variable in the repository's CI configuration file so all pipeline stages can read it.
E.Grant the service principal workspace admin rights so it can deploy any resource without encountering permission errors.
AnswersB, C

A service principal provides non-interactive, auditable automation identity, and injecting its secret from the CI secret manager at runtime keeps credentials out of the repository and logs. This is the recommended way to authenticate deployments in CI/CD. It satisfies the security requirement while allowing the pipeline to deploy to production reliably across runs.

Why this answer

Failing fast and staying secure means validating the bundle before deploying and authenticating with a service principal whose secret comes from the CI secret manager. Validation catches configuration errors before production is touched, and the secret manager keeps credentials out of the repository. Plaintext secrets, admin-level permissions, and branch-agnostic production deploys all increase risk without meeting the stated goals.

Exam trap

The trap here is treating broad permissions or early deployment as safety measures, when least privilege and pre-deploy validation are what actually protect production.

6
MCQmedium

When promoting code from a development workspace to a production workspace, what is the primary risk of using manual notebook exports?

A.The notebooks will run significantly slower in the production environment.
B.The process creates an inconsistent state between development and production.
C.The Databricks workspace will automatically delete the old notebooks.
D.Manual exports are prohibited by Databricks security policies.
AnswerB

Manual processes lack the rigor of automated CI/CD pipelines. Differences in libraries, environment settings, or notebook versions often occur when files are manually copied. These inconsistencies lead to code that works in development but fails in production, causing difficult-to-debug errors and potential downtime for critical data tasks.

Why this answer

Manual notebook exports lack version integrity and metadata synchronization, leading to 'code drift' where the production code doesn't match the source of truth in Git. This manual process is prone to human error, such as importing the wrong version or missing library dependencies. Standardizing on Git-integrated workflows or Databricks Asset Bundles mitigates this risk by automating the deployment and ensuring that exactly what was tested in development reaches production.

Exam trap

Candidates often view manual notebook exports as harmless, overlooking how easily they introduce code drift and lack metadata synchronization across environments.

7
MCQeasy

What is the primary purpose of a 'feature branch' in a Git-based workflow for Databricks?

A.To allow multiple users to edit the same file simultaneously without sync.
B.To isolate changes so they can be tested before merging to production.
C.To bypass the authentication requirements for the workspace.
D.To automatically deploy changes to the production environment.
AnswerB

Branching provides a safe sandbox for development. By working on a feature branch, developers can test their code in a dedicated workspace environment without endangering the current production code. This separation is fundamental to CI/CD, as it ensures that only verified, stable code is eventually merged into the main line.

Why this answer

Feature branches allow developers to build and test new functionality in isolation, without affecting the main production code. This workflow is essential for CI/CD because it facilitates peer reviews, automated integration testing, and safe integration. By keeping the main branch stable and only merging validated features, teams maintain a high standard of quality and avoid breaking the production environment with incomplete or untested code changes, which is a core benefit of modern DevOps practices.

Exam trap

Candidates often confuse feature branches with production deployment branches, missing the core purpose of isolation for testing and peer review prior to merging into the main branch.

8
Multi-Selecthard

A data engineering team is setting up a CI/CD pipeline for Databricks notebooks using Databricks Repos and a Git provider. They want to ensure that changes are tested before being merged and that production deployments are controlled. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Store production credentials in the Repo's notebooks as widgets so the CI job can read them during tests.
B.Run automated tests in a CI job on a feature branch using a dedicated test workspace before merging to main.
C.Merge changes only after the CI pipeline passes and use a separate deployment job to update the production Repo to a tagged release.
D.Allow developers to commit directly to the main branch in the production workspace to reduce merge conflicts.
E.Configure the production Repo to track the main branch and enable automatic pull on every commit.
AnswersB, C

Running tests on feature branches in an isolated test workspace validates changes without affecting production data or users. This practice catches integration issues early and aligns with the CI principle of fast feedback. Using a dedicated workspace also prevents accidental writes to production catalogs and keeps test credentials separate from production credentials.

Why this answer

A robust CI/CD process for Databricks Repos runs automated tests on feature branches in a dedicated test workspace, then gates merges on CI success. Production deployment should be a separate, controlled step that updates the production Repo to a tagged release, providing an immutable and rollback-friendly reference. These two practices together enforce quality gates and controlled promotion, which are core to CI/CD with Databricks Repos.

Exam trap

The trap here is confusing continuous synchronization of the main branch with proper CI/CD, when production should instead be updated through a controlled release step.

9
MCQmedium

A team is designing a CI/CD pipeline using Databricks Asset Bundles (DABs). Which TWO of the following are primary benefits of using DABs for managing Databricks projects?

A.DABs automatically provision cloud-native IAM roles for all users.
B.DABs enable consistent deployment of jobs and pipelines across different environments.
C.DABs provide built-in visual drag-and-drop tools for building ETL pipelines.
D.DABs simplify version control for infrastructure and job configurations.
E.DABs bypass the need for any authentication tokens during deployment.
AnswerB, D

DABs use a declarative YAML structure that allows developers to define jobs, pipelines, and workspace objects once and deploy them consistently across multiple workspaces. This ensures that development, staging, and production environments remain synchronized, preventing the 'works on my machine' syndrome that often plagues manual configuration efforts.

Why this answer

Databricks Asset Bundles represent the modern standard for orchestrating Databricks projects as code. By codifying infrastructure and application configuration into YAML, teams achieve environment parity across development, staging, and production. This approach reduces manual configuration errors, simplifies deployment automation via CLI-based workflows, and provides a version-controlled foundation for managing complex data assets, which is critical for large-scale enterprise data engineering environments.

Exam trap

Candidates often confuse DABs with older manual deployment methods or generic Terraform scripts, failing to recognize that DABs are specifically optimized for Databricks project orchestration, environment parity, and CLI-based automation.

10
MCQmedium

Refer to the exhibit. A CI/CD pipeline running a Databricks CLI command fails with the error shown. What is the most likely cause?

A.The Databricks CLI is not installed on the runner.
B.The path provided in the deployment script does not exist in the target Databricks workspace.
C.The Personal Access Token has expired.
D.The user does not have permissions to write to the root directory.
AnswerB

The API response explicitly states that the path does not exist. In CI/CD pipelines, this usually means the script is referencing a path structure that is inconsistent between environments or that the required folder structure has not been provisioned in the target environment before the deployment script runs.

Why this answer

This error indicates that the automation script is attempting to reference a file or directory path that is missing in the target workspace. In CI/CD, this frequently happens when paths are hardcoded for local development environments rather than being abstracted. Using relative paths or workspace-agnostic configurations within Databricks Asset Bundles prevents these environment-specific failures, ensuring that the deployment process remains robust regardless of where the code is being executed during the pipeline run.

Exam trap

Candidates mistakenly assume the error stems from syntax issues or invalid credentials, overlooking the fact that absolute paths often do not match the remote workspace structure during automation runs.

11
MCQmedium

A data platform team is migrating their deployment process to Databricks Asset Bundles (DABs). They already have a Python wheel task defined in a Databricks Job and a set of notebooks in a Git repository. They want the bundle deployment to be repeatable across development, staging, and production targets with different cluster sizes. Which approach should they take to parameterize the target-specific cluster configuration?

A.Store cluster configuration in a Delta table and have the job read it at runtime using spark.conf.
B.Define cluster settings directly in the job YAML and use separate branches for each environment.
C.Create a separate Databricks Job for each environment and deploy them with the same bundle.
D.Use the bundle's databricks.yml targets block with variables and override cluster node types per target.
AnswerD

DABs support a targets section in databricks.yml where each target can override variables, workspace host, and resource properties. By defining variables for cluster node types and cluster sizes, the team can keep one job definition and supply target-specific values. This is the documented pattern for environment parity and repeatable deployments across dev, staging, and prod.

Why this answer

Databricks Asset Bundles provide a targets block in databricks.yml specifically so one bundle can deploy to multiple workspaces with environment-specific overrides. Using variables for cluster properties and overriding them per target preserves a single job definition while still allowing dev, staging, and prod to use different cluster sizes. This matches the recommended DABs workflow for environment parity and repeatable CI/CD deployments.

Exam trap

The trap here is assuming environment differences require separate job definitions or branches rather than target-level variable overrides within a single bundle.

12
MCQmedium

A data engineering team stores all production notebooks and job definitions in a Git repository. They want every merge to the `main` branch to automatically deploy the updated notebooks to the production Databricks workspace without any manual copy/paste. Which approach should they implement?

A.Schedule a Databricks Job that executes `git pull` inside a notebook on a fixed cron schedule.
B.Enable automatic Git synchronization on the production workspace so Databricks pulls changes from the repository on its own.
C.Use the Databricks CLI to export notebooks from the development workspace and import them into production on every merge.
D.Configure a Databricks Repo in the production workspace that tracks the `main` branch, and run a CI/CD pipeline that calls the Repos API to update the repo to the latest commit after each merge.
AnswerD

A Repo tracking `main` plus a pipeline step that calls the Repos API to pull the latest commit gives an automated, reproducible deployment of notebooks and job definitions. This keeps the production workspace in sync with the source-controlled branch and removes manual export/import steps.

Why this answer

Deploying from Git requires the production workspace to consume the repository as the source of truth. A Databricks Repo that tracks the `main` branch, updated by a CI/CD pipeline through the Repos API after each merge, achieves automated, traceable deployment of notebooks and job definitions without manual copying.

Exam trap

The trap here is assuming Databricks can automatically pull Git changes into a workspace without a pipeline or API call orchestrating the update.

13
MCQeasy

Which component of Databricks CI/CD is responsible for executing automated tests on code before it is merged into the main branch?

A.The Databricks Delta Lake storage layer.
B.The CI/CD runner (e.g., GitHub Actions, Jenkins).
C.The Unity Catalog governance framework.
D.The Databricks SQL Warehouse.
AnswerB

CI/CD runners facilitate the automation of tasks such as running unit tests, linting code, and triggering workspace API calls. They serve as the orchestration layer that verifies the code's integrity in a clean environment, ensuring that any issues are detected before the code is finalized in the repository.

Why this answer

CI/CD runners or build servers (like GitHub Actions, GitLab CI, or Jenkins) are the components responsible for triggering automated testing. By running tests in an isolated environment during the pull request process, these tools ensure that only validated code is merged. This 'fail-fast' approach is critical for maintaining code quality, reducing regression risks, and providing developers with immediate feedback on their changes before they ever impact the production pipeline.

Exam trap

Candidates mistakenly choose Databricks-internal components like the workspace or job scheduler, failing to realize that the external CI/CD runner is the entity responsible for orchestrating the testing workflow.

14
MCQhard

A team is implementing a CI/CD process for their Delta Live Tables (DLT) pipelines. Which THREE of the following practices are recommended to ensure reliable deployment?

A.Keep all DLT configuration values hardcoded within the pipeline notebook.
B.Use Git branches to manage feature development and production code.
C.Implement automated tests that run against the pipeline before production deployment.
D.Manually update the pipeline source code in the production workspace.
E.Define pipeline infrastructure using declarative files like JSON or YAML.
AnswerB, C, E

Branching allows developers to isolate their changes and perform testing without affecting the stable production version. Merging through pull requests ensures that all code changes undergo peer review, which is a critical gatekeeping mechanism for maintaining pipeline stability and preventing unauthorized or faulty code deployments.

Why this answer

Reliable DLT deployments depend on version control, automated testing, and environment abstraction. By defining DLT pipelines as code via DABs or Terraform, teams ensure consistent environments. Testing code before deployment prevents runtime errors, and using Git-based workflows allows for code reviews, which are essential for maintaining high-quality pipelines.

These practices collectively ensure that DLT pipeline changes are predictable, verifiable, and safe to deploy into production environments without disrupting existing operations.

Exam trap

Candidates often assume that manual testing in the UI is sufficient, overlooking the requirement for automated, repeatable tests and declarative infrastructure definitions that characterize robust CI/CD pipelines.

15
MCQhard

Refer to the exhibit. A CI/CD pipeline fails with the provided error. What is the most likely cause?

A.The branch 'main' does not exist in the remote repository.
B.The Service Principal lacks sufficient permissions for the workspace path.
C.The Databricks CLI is not installed on the runner.
D.The Git repository URL is invalid.
AnswerB

A 403 Forbidden status code confirms that the authentication was successful, but the authorized identity is not permitted to perform the specific operation. This is a common permission-management issue in CI/CD where the automation identity hasn't been granted access to the specific Repos folder in the Databricks workspace.

Why this answer

The 403 Forbidden error indicates that the identity running the command lacks the necessary permissions to perform the action. In a CI/CD context, this usually happens because the Service Principal lacks 'Can Manage' or 'Can Edit' permissions on the target Repos folder. Ensuring the Service Principal has the correct workspace-level permissions is a standard requirement for successful automated deployments, preventing access issues during the sync process.

Exam trap

Candidates often guess that the error is due to network connectivity or syntax issues, overlooking that a 403 error in automated deployments is almost always a Service Principal permission deficiency.

16
MCQeasy

A team wants to automate deployment of Databricks jobs from a Git repository using a CI/CD pipeline. They need a tool that reads a declarative project definition and creates or updates jobs, pipelines, and notebooks in a target workspace. Which Databricks capability should they use?

A.Cluster policies, by attaching a policy that references the Git repository and applies job configurations at cluster start.
B.Databricks Asset Bundles, using databricks bundle deploy to apply the project definition to the target workspace.
C.Databricks Repos, by adding the repository and clicking Sync in the target workspace before each release.
D.The Databricks SQL warehouse API, using scheduled queries to create jobs and pipelines in the target workspace.
AnswerB

Databricks Asset Bundles let you define jobs, pipelines, and notebooks declaratively in databricks.yml along with the source files. Running databricks bundle deploy reads that definition and creates or updates the corresponding resources in the target workspace. This is the supported mechanism for automated, repeatable deployment from a Git repository in a CI/CD pipeline.

Why this answer

Databricks Asset Bundles are the declarative deployment mechanism for Databricks resources. A databricks.yml file plus source files defines jobs, pipelines, and notebooks, and databricks bundle deploy applies that definition to a target workspace. Repos handle file sync for development, SQL warehouse APIs run queries, and cluster policies govern compute; none of them deploys resource definitions from Git.

Exam trap

The trap here is confusing Databricks Repos, which syncs files for interactive work, with Asset Bundles, which deploy resource definitions declaratively.

17
Multi-Selectmedium

Which TWO of the following statements are correct regarding the use of Databricks Repos for CI/CD?

Select 2 answers
A.Databricks Repos allows users to integrate with Git providers like GitHub, GitLab, and Bitbucket.
B.Databricks Repos can only be used for Python code and does not support SQL or Markdown files.
C.Each Databricks Repo is a local-only folder that cannot be synchronized with a remote Git repository.
D.Databricks Repos allows for managing multi-branch workflows and pull requests through Git.
E.Changes made in a Repo automatically deploy to production without needing to be merged in Git.
AnswersA, D

Databricks Repos natively supports major Git providers, including GitHub, GitLab, and Bitbucket. This allows teams to clone, pull, push, and manage branches directly from the Databricks UI or API, fostering a unified workflow that keeps data science and engineering code synchronized with standard enterprise source control systems.

Why this answer

Databricks Repos facilitates seamless Git integration within the Databricks workspace. It enables teams to perform development directly in the workspace while keeping code versioned in remote Git repositories. By understanding that Repos maps workspace folders to Git branches, engineers can implement effective branching strategies and merge workflows.

This integration bridges the gap between collaborative notebook development and standard software engineering practices, ensuring that changes are tracked, reviewed, and tested before deployment to production environments.

Exam trap

Candidates frequently assume Databricks Repos replaces external CI/CD runners or limits version control capabilities, forgetting that it natively integrates with popular Git providers and supports robust multi-branch workflows.

18
MCQmedium

A data engineering team stores its Databricks notebooks and Python files in a Git repository. They want to avoid manually copying files into the workspace and ensure that the production workspace always runs the exact code version that passed tests. Which approach should they use?

A.Schedule a Databricks Job that runs a notebook which pulls the latest code from Git at runtime.
B.Mount the Git repository as a DBFS mount point and reference notebooks directly from the mount.
C.Use Databricks Repos to clone the Git repository into the production workspace and check out the tested commit.
D.Export notebooks as .dbc files and import them into the production workspace using the Databricks CLI.
AnswerC

Databricks Repos integrates directly with Git, allowing the workspace to clone a repository and check out a specific commit or tag. This ensures the production workspace executes exactly the code version that passed tests, without manual file copying. It also provides version traceability and supports CI/CD automation through the Repos API.

Why this answer

Databricks Repos provides native Git integration, enabling teams to clone repositories and check out specific commits. By checking out the tested commit in the production workspace, the team ensures that the deployed code matches what passed tests, eliminating manual copying and reducing drift. This is a core practice for implementing CI/CD with Databricks.

Exam trap

The trap here is assuming that any file transfer method (like .dbc export or DBFS mount) is equivalent to Git-based deployment, when only Databricks Repos provides commit-level traceability and native execution.

19
MCQmedium

A data engineering team stores notebooks in a Git repository and wants automated deployments to a Databricks workspace. They configure a GitHub Actions workflow that runs a Databricks CLI command to deploy Databricks Asset Bundles. The workflow authenticates using a service principal OAuth token stored in GitHub Secrets. After the first successful run, subsequent runs fail with an authentication error. The token was created with a 1-hour lifetime. What should the team do to ensure the workflow can authenticate reliably on every run?

A.Run the GitHub Actions workflow on a self-hosted runner inside the Databricks workspace network so that the CLI can use an instance profile for authentication.
B.Increase the service principal OAuth token lifetime to 24 hours in the Databricks account console and update the secret in GitHub.
C.Configure the GitHub Actions workflow to obtain a short-lived OAuth token dynamically from the identity provider using a federated identity or client credentials flow before each deployment.
D.Replace the service principal OAuth token with a personal access token (PAT) generated for a workspace admin user and store it in GitHub Secrets.
AnswerC

Service principal OAuth tokens are short-lived. To authenticate reliably on every run, the pipeline must request a fresh token at runtime rather than reuse a static secret. Using a client credentials flow or workload identity federation with the identity provider lets the workflow mint a token per run, avoiding expiration failures and manual rotation.

Why this answer

Short-lived OAuth tokens expire, so a CI/CD pipeline must fetch a new token at the start of each run. Using the client credentials flow or workload identity federation with the service principal allows the GitHub Actions workflow to authenticate dynamically without storing a long-lived secret. This is the secure, maintainable pattern for automated Databricks deployments.

Exam trap

The trap here is assuming that any token stored in GitHub Secrets will remain valid indefinitely, when short-lived service principal OAuth tokens expire and must be regenerated per run.

20
MCQmedium

A team is using Databricks Asset Bundles (DABs) to manage their CI/CD pipeline. They want to run unit tests on their Python code before deploying the bundle. Where should the tests be executed in the pipeline?

A.Using the Databricks CLI to run a test command that executes tests on a cluster.
B.Inside a Databricks notebook that is executed manually by a developer before merging.
C.In the CI/CD runner (e.g., GitHub Actions) as a separate job step before the bundle deploy command.
D.As a Databricks job within the bundle that is triggered after deployment.
AnswerC

Running unit tests in the CI/CD runner before deployment ensures that code is validated in isolation, without affecting the Databricks workspace. This is a standard CI practice: test first, then deploy only if tests pass. It also allows the use of standard Python testing frameworks like pytest.

Why this answer

Unit tests should be executed in the CI/CD runner as a step before deployment. This ensures that only code that passes tests is deployed. It also leverages the runner's environment and standard testing tools, keeping the pipeline efficient and reliable.

Deploying first would risk introducing bugs into production.

Exam trap

The trap here is thinking that tests must run inside Databricks, when in fact unit tests are best run in the CI environment before any deployment occurs.

21
MCQeasy

Why is it important to use Service Principals instead of Personal Access Tokens (PATs) for CI/CD automation in Databricks?

A.Service Principals are faster to execute than personal tokens.
B.Service Principals do not require any permission configuration.
C.Service Principals are not tied to an individual's lifecycle.
D.Service Principals allow anyone in the organization to run pipelines.
AnswerC

Service Principals are machine identities that persist regardless of individual employee status. This prevents the common issue where automated pipelines break due to the expiration or revocation of a user's token. They provide a stable, manageable foundation for long-term CI/CD automation that aligns with enterprise security policies.

Why this answer

Service Principals provide a secure and stable identity for automated systems, independent of individual user accounts. Because PATs are tied to a specific user, they become invalid when that user leaves the company or their access is revoked. Service Principals ensure continuous operation of CI/CD pipelines, simplify security auditing, and follow the principle of least privilege, which is crucial for enterprise-grade automation and governance.

Exam trap

Candidates mistakenly believe Personal Access Tokens are secure enough for production CI/CD, overlooking the critical risk of user offboarding breaking pipelines.

22
MCQmedium

Which of the following is a recommended strategy for managing library dependencies in a CI/CD pipeline for Databricks?

A.Install libraries manually on the cluster via the Databricks UI.
B.Include a requirements.txt file and install libraries as part of the job definition.
C.Use the latest version of all libraries to ensure best performance.
D.Store library binary files directly in the Git repository.
AnswerB

Using a requirements file allows for version-controlled dependency management. Defining these libraries within the job configuration ensures that every time a job is triggered or updated by the CI/CD pipeline, the necessary packages are installed, ensuring environment consistency and enabling reproducible results across different workspaces and development stages.

Why this answer

Managing libraries through a centralized repository or a requirements.txt file ensures that the exact same versions are installed consistently across all environments. This avoids the 'dependency hell' where code fails because of incompatible library versions. By integrating dependency management into the CI/CD build process, teams can audit used versions, ensure security compliance, and guarantee that the production pipeline runs with the exact dependencies that were tested during the development phase.

Exam trap

Candidates often suggest installing libraries manually in the UI or via init scripts, failing to realize that requirements.txt integration ensures consistency and auditability throughout the CI/CD lifecycle.

23
Multi-Selecthard

A data engineering team is implementing a CI/CD pipeline for Databricks notebooks and jobs using Databricks Asset Bundles. They want to ensure deployments are reproducible and that production changes are traceable. Which TWO of the following practices should they follow? (Choose two.)

Select 2 answers
A.Deploy the bundle from a developer's local machine to production to avoid the complexity of a CI/CD runner.
B.Use a separate target in databricks.yml for each environment, and run databricks bundle deploy with the appropriate target in the pipeline.
C.Manually edit the deployed job in the production workspace after deployment to apply urgent fixes, and sync those changes back to Git later.
D.Embed production credentials directly in the databricks.yml file so the pipeline can authenticate without additional configuration.
E.Store the databricks.yml bundle configuration and all referenced notebooks in the same Git repository, and deploy from a specific Git commit SHA.
AnswersB, E

Targets allow a single bundle to be deployed to different workspaces with environment-specific settings, such as workspace host and variable values. Invoking the deploy command with the correct target in each pipeline stage ensures consistent, repeatable deployments. This is the recommended way to manage multiple environments with Databricks Asset Bundles.

Why this answer

Reproducible and traceable deployments require that the bundle and notebooks are versioned together and deployed from a specific commit. Using environment-specific targets in databricks.yml ensures consistent deployment across workspaces. Manual edits, embedded credentials, and local deployments all break reproducibility and traceability, so they must be avoided.

Exam trap

The trap here is thinking that syncing manual production edits back to Git later is acceptable, when any manual edit breaks the link between the deployed state and a specific commit.

24
MCQmedium

What is the primary goal of implementing 'environment parity' in a Databricks CI/CD pipeline?

A.To ensure every developer has full admin access to production.
B.To minimize failures caused by differences between testing and production.
C.To reduce the cost of compute resources in development workspaces.
D.To allow for manual code deployments in all environments.
AnswerB

The biggest risk to a CI/CD process is the 'it worked in staging but failed in production' scenario. Parity ensures that the runtime environment is consistent across stages, allowing for reliable testing. This predictability is vital for high-quality data engineering, as it ensures that pipeline logic is thoroughly validated.

Why this answer

Environment parity ensures that the development, staging, and production environments are identical in configuration, libraries, and settings. This eliminates 'environment drift,' where code fails in production due to subtle discrepancies. Achieving parity through Infrastructure-as-Code (IaC) is the cornerstone of reliable CI/CD, as it guarantees that testing in staging provides a true representation of how the code will behave in production, significantly reducing the frequency of deployment-related failures.

Exam trap

Candidates often mistake 'environment parity' for simply having the same data, failing to realize it specifically refers to identical infrastructure, library, and configuration settings across environments.

25
MCQeasy

A team is using the Databricks CLI in their CI/CD pipeline to deploy jobs and notebooks. They need the pipeline to authenticate to a production workspace without embedding a personal user's credentials. Which authentication method should they configure for the CLI?

A.A personal access token generated from a developer's Databricks account.
B.A service principal with an OAuth token or client credentials configured in the CLI profile.
C.The workspace's built-in admin username and password stored in the CI environment variables.
D.An SSH key pair registered in the Databricks workspace under the deploying user's profile.
AnswerB

Service principals are non-human identities designed for automation. The Databricks CLI supports authenticating as a service principal using OAuth client credentials or a service principal token, which can be scoped to only the needed workspace permissions. This avoids dependency on individual users and supports rotation and auditing in CI/CD.

Why this answer

For CI/CD automation, Databricks recommends authenticating the CLI with a service principal rather than a personal user identity. Service principals are non-human accounts that can be granted scoped permissions, and the CLI supports OAuth client credentials or service principal tokens. This avoids coupling deployments to an employee's account and supports credential rotation, auditing, and least privilege in production pipelines.

Exam trap

The trap here is assuming any valid Databricks credential works for CI/CD, when automation should use a non-human service principal rather than a personal token.

26
MCQmedium

A data engineering team wants to implement Git integration for their Databricks notebooks. Which workflow is considered the best practice for CI/CD in Databricks Repos?

A.Export notebooks manually to local machines and commit them via Git CLI.
B.Use the Databricks REST API to push code updates directly to production notebooks.
C.Develop code in feature branches, merge via pull requests, and use Databricks Repos to sync production.
D.Develop code in the production workspace and use Git only for final snapshots.
AnswerC

This workflow leverages standard DevOps practices like branching and pull requests to ensure code quality through peer reviews. By using Databricks Repos to sync, the production environment stays consistent with the verified main branch, minimizing configuration errors and ensuring that only tested, approved code is executed in production workflows.

Why this answer

Integrating Databricks Repos with a Git provider like GitHub enables version control at the notebook level. This workflow allows teams to use feature branches for development, submit pull requests for code review, and merge into a main branch that triggers automated deployments. Adopting this standardizes the development lifecycle, ensures code traceability, and prevents manual, error-prone deployments in production environments, which is essential for maintaining production-grade data pipelines.

Exam trap

Candidates often suggest manual notebook exports or direct Git commits to production branches, ignoring the necessity of pull requests for code review and automated deployment workflows.

27
MCQmedium

A data engineer is setting up a CI/CD pipeline that runs unit tests on transformation logic before deploying notebooks to a production Databricks workspace. The tests must run quickly and not depend on a live Databricks cluster or external data sources. Which approach best meets these requirements?

A.Extract transformation logic into pure Python functions and test them locally with a framework like pytest, using sample DataFrames created in memory.
B.Run the tests as a Databricks job on a small interactive cluster in the development workspace, using production data for realism.
C.Configure the pipeline to trigger a Delta Live Tables pipeline in the development workspace and check the pipeline's event log for errors.
D.Use Databricks Repos to run notebooks interactively and manually verify the output before merging the pull request.
AnswerA

Extracting logic into pure Python functions allows unit tests to run locally without a cluster or external data. Using pytest with in-memory DataFrames (for example, via PySpark local mode or pandas) keeps tests fast and deterministic. This is the standard approach for testing transformation logic in a CI/CD pipeline before deployment.

Why this answer

Unit tests for transformation logic should be fast, isolated, and runnable without a Databricks cluster. Extracting logic into pure Python functions and testing them with pytest and in-memory DataFrames achieves this. This practice also encourages modular code that is easier to maintain and integrate into CI/CD pipelines.

Exam trap

The trap here is equating any automated test on Databricks with a unit test, when true unit tests should avoid cluster dependencies and external data entirely.

28
MCQhard

A data engineering manager wants to ensure that all production code in Databricks is fully audited and versioned. Which TWO of the following steps are required?

A.Grant developers 'Can Manage' permissions on production workspace folders.
B.Enforce a policy where only the CI/CD Service Principal can update production assets.
C.Store all notebook development directly in the production workspace.
D.Enable Git integration and require pull requests for all merges into main.
E.Allow administrators to use personal tokens for manual emergency hotfixes.
AnswerB, D

This policy ensures that human users cannot bypass the CI/CD pipeline to make unauthorized or un-audited changes. By restricting updates to a machine identity, the organization guarantees that every change is captured, reviewed, and deployed according to the defined CI/CD process, which is critical for audit compliance and environment stability.

Why this answer

Ensuring auditability and versioning requires locking down the production environment and forcing all changes through a Git-based CI/CD pipeline. By disabling manual changes in production and requiring all updates to be merged through a pull request process, teams ensure that every modification is reviewed, logged, and linked to a specific version. These controls are essential for compliance, stability, and maintaining a clear history of system changes in professional data environments.

Exam trap

Candidates often suggest manual reviews or periodic audits, failing to grasp that true auditability requires programmatic enforcement via Git-based workflows and restricted Service Principal access to production environments.

29
MCQmedium

A team uses Databricks Asset Bundles to define a job that must exist in both a staging and a production workspace with different cluster sizes. They want a single bundle definition that deploys correctly to both targets. Which configuration should they use?

A.Define two separate bundles, one per workspace, and deploy each with its own `databricks bundle deploy` command.
B.Use a single bundle with `targets` for staging and production, and override cluster settings per target using target-specific variables or overrides in `databricks.yml`.
C.Hard-code the production cluster size in the bundle and use a cluster policy in staging to downsize it at runtime.
D.Deploy the bundle only to production and point the staging workspace at the production job through a shared workspace link.
AnswerB

Databricks Asset Bundles allow multiple `targets` in one `databricks.yml`, each with its own workspace host and overrides. Cluster sizes can differ per target through target-scoped variables or resource overrides, so one bundle definition deploys correctly to both staging and production.

Why this answer

Databricks Asset Bundles are designed for multi-environment deployment. A single `databricks.yml` with `targets` for staging and production, combined with target-specific overrides for cluster configuration, lets one bundle definition produce environment-appropriate resources without duplicating the job definition.

Exam trap

The trap here is assuming cluster policies or shared workspace links can substitute for target-specific bundle overrides.

30
MCQmedium

Your team is migrating a manual Databricks job to a CI/CD pipeline. You need to ensure the job configuration is version-controlled and deployed programmatically. Which approach aligns with Databricks best practices?

A.Manually export the job JSON via the UI and copy-paste it into the production workspace.
B.Use the Databricks UI to update production jobs directly to ensure changes are applied immediately.
C.Define the job configuration in YAML files using Databricks Asset Bundles and deploy via the CLI.
D.Write a custom Python script that uses the workspace API to recreate the job from scratch every time.
AnswerC

Databricks Asset Bundles allow you to define jobs as code in YAML, which can be stored in Git. Using the Databricks CLI to deploy these files enables automated, consistent, and repeatable deployments across different environments, which is the cornerstone of robust CI/CD practices in the Databricks platform.

Why this answer

Using Databricks Asset Bundles (DABs) is the recommended approach for modern CI/CD. It treats infrastructure as code, allowing you to define jobs, pipelines, and workflows in YAML files. This ensures consistency across environments like development, staging, and production.

By decoupling the job definition from the workspace UI, you gain auditability, reproducibility, and the ability to integrate seamlessly with Git workflows, reducing manual configuration errors and accelerating release cycles for data engineering teams.

Exam trap

Candidates often choose manual UI deployment or legacy JSON templates instead of modern Databricks Asset Bundles (DABs) combined with YAML configuration files, which are the current industry best practice for programmatic CI/CD pipelines.

31
MCQhard

A team is using Databricks Asset Bundles (DABs) to deploy a job to multiple environments (dev, staging, prod). They need to ensure that the job uses different cluster sizes and schedules per environment. Which DABs feature should they use?

A.Use the Databricks REST API to update the job configuration after deployment with environment-specific values.
B.Create a separate Databricks workspace for each environment and hardcode the cluster size and schedule in the job definition.
C.Define separate bundle configuration files for each environment and use the --target flag with databricks bundle deploy.
D.Use Jinja2 templating in the job definition to inject environment-specific values from environment variables.
AnswerC

Databricks Asset Bundles support targets, which are defined in the databricks.yml file. Each target can override resource properties such as cluster size and schedule. Using the --target flag selects the appropriate configuration for deployment. This is the intended way to manage environment-specific settings in DABs.

Why this answer

Databricks Asset Bundles use targets to define environment-specific overrides in the databricks.yml file. When deploying, the --target flag selects the target, applying the corresponding cluster size and schedule. This keeps a single source of truth while allowing differences per environment, which is essential for promoting bundles through dev, staging, and prod.

Exam trap

The trap here is thinking that external templating or post-deployment API calls are needed, when DABs natively support environment-specific configuration through targets.

32
MCQmedium

A data engineer is configuring a CI/CD pipeline for Databricks notebooks using GitHub Actions. They need to authenticate to Databricks to deploy notebooks. Which authentication method is recommended for production CI/CD pipelines?

A.Use the Databricks CLI with interactive login during the CI run.
B.Store a personal access token (PAT) in GitHub Secrets and use it in the workflow.
C.Use a service principal with OAuth tokens generated via the Databricks CLI or REST API, storing the client ID and secret in GitHub Secrets.
D.Embed the Databricks username and password directly in the workflow YAML file.
AnswerC

Service principals are non-human identities that can be granted specific permissions. Using OAuth tokens with a client ID and secret stored as secrets in GitHub Actions allows secure, automated authentication. This is the recommended practice for production CI/CD because it avoids personal credentials and supports fine-grained access control.

Why this answer

For production CI/CD, authentication should use a service principal with OAuth. The client ID and secret are stored as secrets in the CI system, such as GitHub Secrets. The workflow can then use the Databricks CLI or REST API to authenticate.

This approach decouples automation from individual users, supports permission scoping, and is more secure than personal access tokens.

Exam trap

The trap here is assuming that a personal access token is sufficient for CI/CD, when service principals with OAuth are the recommended method for production.

33
MCQmedium

A team stores Databricks notebooks and job definitions in a Git repository and wants every merge to the main branch to automatically deploy to production. Which combination of practices should the pipeline implement to achieve this safely?

A.Trigger the deployment workflow on a nightly schedule and use a personal access token for authentication so the deployment uses the developer's existing permissions.
B.Trigger the deployment workflow on pushes to any branch and use the Databricks CLI with interactive browser login to authenticate to the production workspace.
C.Trigger the deployment workflow on pull requests and use the developer's personal access token stored in the repository as a plaintext secret.
D.Trigger the deployment workflow on pushes to the main branch, authenticate with a service principal whose credentials are stored in the CI system's secret manager, and run databricks bundle deploy.
AnswerD

Deploying on pushes to main ensures only merged code reaches production, and a service principal stored in the CI secret manager provides non-interactive, auditable authentication. Running databricks bundle deploy applies the bundle definition from the merged commit. This combination enforces the review gate, keeps credentials out of the repository, and makes deployments reproducible and traceable.

Why this answer

Safe continuous deployment requires two things: a trigger scoped to the main branch so only reviewed, merged code reaches production, and non-interactive authentication via a service principal stored in the CI secret manager. Running databricks bundle deploy from the merged commit then applies the bundle definition reproducibly. Scheduling, branch-agnostic triggers, and personal tokens all undermine control or automation.

Exam trap

The trap here is focusing on the authentication method alone and overlooking that the workflow trigger must be scoped to the main branch to prevent unreviewed code from reaching production.

34
MCQhard

When utilizing Databricks Asset Bundles, how should secrets (e.g., API keys, database credentials) be handled to ensure security during CI/CD?

A.Hardcode the secrets in the bundle's YAML configuration files.
B.Store secrets as environment variables in the Git repository.
C.Use the bundle's secret reference syntax to map variables to a secret scope.
D.Ask all developers to manually enter secrets in the production UI.
AnswerC

Referencing secret scopes allows the CI/CD pipeline to inject credentials at execution time without exposing the actual values in the code. This is the industry-standard approach to managing sensitive information, ensuring that credentials remain secure while allowing the automation engine to authenticate successfully to external services during job execution.

Why this answer

Hardcoding secrets in configuration files is a critical security vulnerability. Databricks Asset Bundles support referencing secrets through the 'secrets' scope, which securely retrieves credentials at runtime. By decoupling sensitive data from the version-controlled YAML files, teams maintain security, avoid leaking credentials in Git, and simplify secret rotation.

This practice is essential for enterprise security and ensures that sensitive infrastructure data is managed in alignment with the principle of least privilege.

Exam trap

Candidates frequently suggest storing secrets in environment variables or configuration files directly, ignoring the security risk of leaking credentials and the availability of Databricks' built-in secret reference syntax.

35
MCQmedium

Why should developers avoid using 'notebook' references in production pipelines that point to the 'Shared' folder for shared development work?

A.Databricks jobs cannot execute notebooks from the 'Shared' directory.
B.The 'Shared' directory is deleted automatically every 24 hours.
C.It lacks the version control and access isolation required for production.
D.Notebooks in the 'Shared' directory are limited to a smaller file size.
AnswerC

Production assets require strict access controls and a clear audit trail. The 'Shared' folder usually has loose permissions, allowing any user to edit the code. This creates a high risk of production instability. Proper CI/CD processes dictate that production assets are deployed to hardened, versioned folders via automated pipelines.

Why this answer

The 'Shared' folder is meant for collaborative development, not for hosting production-ready code. Pipelines referencing this folder are susceptible to unpredictable changes, as any developer with access can modify the code. In CI/CD, production code should reside in a controlled, versioned, and immutable location.

Using the 'Shared' folder for production pipelines violates the separation of concerns and increases the risk of unauthorized or accidental changes breaking critical data workflows.

Exam trap

Candidates often underestimate the risks of the 'Shared' folder, incorrectly believing it is a valid location for production code if permissions are restricted, ignoring the need for immutable version control.

36
MCQeasy

When integrating Databricks with a CI/CD tool like GitHub Actions, how should you securely manage the authentication token used for deployments?

A.Hardcode the Databricks Personal Access Token directly into the YAML pipeline file.
B.Store the token as an encrypted secret in the CI/CD provider's secret store.
C.Save the token in a public S3 bucket and download it when the pipeline starts.
D.Use your own password as the token for easier memorization and access.
AnswerB

Storing credentials in a secure, encrypted secret store provided by CI/CD platforms like GitHub Actions or GitLab is the standard security practice. This prevents exposure in logs or source control while ensuring the pipeline can securely access the Databricks API during the execution of deployment tasks.

Why this answer

Managing secrets securely is vital for CI/CD. Storing tokens directly in code or pipeline definitions exposes them to unauthorized access. By using a secrets manager, you decouple the sensitive credentials from the pipeline logic.

This ensures that the CI/CD platform can authenticate with Databricks dynamically during the execution phase, maintaining a high security posture while allowing for easy rotation of credentials without needing to refactor the entire codebase.

Exam trap

Candidates often suggest hardcoding tokens in the repo or using insecure plain-text files, failing to recognize that CI/CD providers offer dedicated encrypted secret stores for this exact purpose.

37
MCQhard

A CI/CD pipeline must run unit tests on Python transformation code before deploying a Databricks job. The tests should execute quickly without starting a cluster and should validate the transformation logic in isolation. Which approach best meets these requirements?

A.Deploy the job to a development workspace and run an integration test against production data to confirm the transformations produce expected results.
B.Configure the pipeline to run databricks bundle validate and rely on schema validation to confirm the transformation logic is correct.
C.Use pytest in the CI runner to test the transformation functions directly, mocking Spark or using a local SparkSession only where needed.
D.Run the tests as a Databricks job task on an all-purpose cluster so the tests use the same runtime as production.
AnswerC

pytest runs in the CI runner with no cluster, so it is fast and isolated. Transformation logic written as pure functions can be tested directly, and a local SparkSession can be created only for tests that need Spark APIs. This keeps the feedback loop short, avoids cloud compute costs, and validates logic before deployment, which is exactly what a pre-deploy unit test stage should do.

Why this answer

Unit tests for transformation logic belong in the CI runner, not on a cluster. Using pytest with pure functions and a local SparkSession where necessary keeps tests fast, isolated, and free of cloud compute. This catches logic errors before deployment, whereas bundle validation only checks configuration and integration tests against production data are slow and risky.

Exam trap

The trap here is conflating bundle validation with unit testing, assuming a successful databricks bundle validate means the transformation logic has been verified.

38
MCQhard

A team is using Databricks Asset Bundles to manage a job that writes to a Unity Catalog table. The bundle is deployed to a staging workspace for testing and then to a production workspace. The team wants the job to use different catalog and schema names in each environment without duplicating the entire bundle. Which approach should they use?

A.Define bundle variables for the catalog and schema, and set different values for each variable in the staging and production targets within the databricks.yml file.
B.Store the catalog and schema names in a Databricks secret scope and reference them in the job's notebook using dbutils.secrets.get.
C.Use a single target and pass the catalog and schema names as job parameters at runtime through the Databricks Jobs API.
D.Create two separate bundle configuration files, one for staging and one for production, and manually copy the job definition into each.
AnswerA

Databricks Asset Bundles support variables that can be overridden per target. By defining catalog and schema variables and assigning environment-specific values under each target in databricks.yml, the same job definition can deploy to staging and production with the correct Unity Catalog names. This is the intended way to parameterize bundles across environments.

Why this answer

Databricks Asset Bundles allow a single job definition to be deployed to multiple environments by using variables that are resolved differently per target. Defining catalog and schema variables and setting their values under the staging and production targets in databricks.yml keeps the bundle DRY and ensures each workspace uses the correct Unity Catalog objects.

Exam trap

The trap here is thinking that runtime job parameters or secret scopes are the right place for environment-specific catalog names, when bundle variables and targets are designed exactly for this purpose.

39
MCQeasy

A team wants to ensure that code changes to their Databricks notebooks are reviewed before being deployed to production. They use a Git repository and Databricks Repos. Which practice should they implement in their CI/CD process?

A.Allow developers to commit directly to the main branch and deploy automatically to production on every commit.
B.Use Databricks Repos to create a separate repository for each developer and merge changes manually in the production workspace.
C.Require pull requests with at least one approval before merging into the main branch, and trigger production deployment only after the merge.
D.Configure the production workspace to allow only workspace admins to edit notebooks, and have developers request admin assistance for each change.
AnswerC

Branch protection rules that require pull requests and approvals enforce peer review before code reaches main. Triggering production deployment after the merge ensures that only reviewed and approved code is deployed. This is a foundational CI/CD practice for controlled, auditable releases.

Why this answer

Enforcing pull requests with approvals through branch protection rules ensures that changes are reviewed before merging. Deploying to production only after the merge ties deployment to reviewed code. This practice is essential for maintaining code quality and auditability in a CI/CD pipeline.

Exam trap

The trap here is confusing access control (who can edit) with change control (who can approve and merge), when code review requires a pull request workflow.

40
MCQhard

A CI/CD pipeline runs unit tests against transformation logic before deploying to production. The tests must run on a Databricks cluster and produce a pass/fail result that fails the pipeline when assertions do not hold. Which implementation best fits this requirement?

A.Run pytest locally in the CI runner against a sample of production data downloaded from cloud storage.
B.Use the Databricks SQL Statement Execution API to run `SELECT` statements that compare row counts between source and target tables.
C.Configure the cluster's init script to run the test suite automatically whenever the cluster starts.
D.Create a Databricks Job that runs a notebook containing assertions, and have the CI/CD pipeline trigger the job and check its run result.
AnswerD

A Databricks Job that executes assertion notebooks on a real cluster exercises the production runtime and returns a run result the pipeline can inspect. If the job fails, the pipeline fails, giving the required pass/fail gate with the same Spark environment used in production.

Why this answer

Tests must run where the production code runs. Triggering a Databricks Job that executes assertion notebooks lets the CI/CD pipeline inspect the run result and fail on assertion errors. This validates behavior on the actual Databricks runtime rather than in a disconnected local environment.

Exam trap

The trap here is treating local pytest runs or SQL row-count checks as equivalent to executing assertions on the Databricks runtime.

Ready to test yourself?

Try a timed practice session using only Implementing Cicd questions.