Courseiva

PMLE Automating and Orchestrating ML Pipelines Practice Question

A team has a Vertex AI pipeline that includes a container component for data preprocessing. The team notices that the component is re-executed every time the pipeline runs, even when the inputs and code haven't changed. They want to leverage pipeline caching to avoid redundant executions. What should they do to enable caching for this component?

⚠ Common exam trap

Many candidates assume caching must be explicitly enabled (like in some other cloud platforms), but Vertex AI caches by default, so the issue is usually that caching was explicitly disabled on the component.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Ensure that the component does not have 'dsl.cache_options(enable_cache=False)' set.

Vertex AI pipeline caching is enabled by default for all components unless explicitly disabled using `dsl.cache_options(enable_cache=False)`. The component re-executing every time indicates that caching was likely disabled on that specific component. Removing or ensuring this setting is not present will allow the pipeline to reuse cached outputs when inputs and code have not changed.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set the 'caching' flag to 'True' in the pipeline definition using 'pipeline.caching = True'.

    Why it's wrong here

    Vertex AI Pipelines exposes no 'pipeline.caching' attribute; caching is set per task through 'enable_caching=True' on the component or task, so this assignment has no effect. It is tempting because pipeline objects hold configuration attributes, and would be correct for pipeline-level settings such as parameters or labels.

  • ✗

    Set the environment variable 'ENABLE_CACHE' to 'true' on the pipeline run request.

    Why it's wrong here

    No 'ENABLE_CACHE' environment variable governs Vertex AI pipeline caching; the run request cannot toggle it, so the component still re-executes. It is tempting because environment variables configure container behaviour, and would be correct for passing runtime values into a component's code, not for pipeline-level caching.

  • ✗

    Re-compile the pipeline with the '--enable-cache' flag.

    Why it's wrong here

    Vertex AI Pipelines has no '--enable-cache' compilation flag; caching is controlled per-task via the component's 'enable_caching' argument, so recompiling changes nothing. It is tempting because CLI flags commonly toggle pipeline features, and would be correct for settings exposed at compile time, such as pipeline parameters or template paths.

  • ✓

    Ensure that the component does not have 'dsl.cache_options(enable_cache=False)' set.

    Why this is correct

    Caching is enabled by default in Kubeflow Pipelines; the component re-executes because cache_options(enable_cache=False) was explicitly set, disabling it. Removing that setting restores default caching behaviour, so unchanged inputs and code reuse the cached execution instead of rerunning.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A team notices that a Vertex AI Pipeline step re-executes every time the pipeline runs, even though its inputs and code have not changed. They want to enable caching for this component to avoid redundant computation. However, caching is currently disabled globally. Which configuration change will enable caching for that specific component?

hard
  • A.Use the 'enable_caching' parameter when creating the pipeline job via the SDK.
  • ✓ B.Add 'caching=True' to the @dsl.component decorator for that component.
  • C.Add 'caching=True' to the @dsl.pipeline decorator.
  • D.Set the environment variable 'CACHE_ENABLED=True' on the Vertex AI Pipeline job.

Why B: In Vertex AI Pipelines (built on Kubeflow Pipelines v2), caching is controlled at the component level via the `@dsl.component` decorator's `caching` parameter. Setting `caching=True` on the specific component enables the pipeline to reuse the cached output when the component's inputs, code, and environment are unchanged, even if caching is disabled globally. This is the only option that targets a single component rather than the entire pipeline or job.

Variation 2. You are building a Vertex AI pipeline using the KFP SDK v2. One component processes a large dataset and outputs a metrics artifact. You notice that the component is being cached even when the dataset changes, because the component code and image remain the same. How can you force the component to always re-execute when the dataset changes?

hard
  • A.Use the dsl.CacheKey annotation to explicitly set cache keys.
  • ✓ B.Set caching_strategy.max_cache_staleness = "0s" on the component.
  • C.Change the component's image tag to :latest.
  • D.Add a random integer parameter to the component to vary the inputs.

Why B: In KFP SDK v2, setting caching_strategy.max_cache_staleness = "0s" on a component forces the pipeline to treat any cached result as stale, effectively disabling cache reuse and ensuring the component re-executes on every run. This is the documented mechanism to override caching when inputs like a dataset change but the component code and image remain identical.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.