PMLE Automating and Orchestrating ML Pipelines Practice Question
A team runs a Vertex AI pipeline that includes a component which downloads a large dataset from BigQuery and writes it to Cloud Storage. The pipeline's caching is enabled by default. During iterative development, the engineer modifies the SQL query inside the component to include an additional feature column, but the pipeline still uses the previously cached output because the component's input parameters and code hash are unchanged. The engineer needs the component to re-execute with the updated query without disabling caching for the entire pipeline. What should the engineer do?
⚠ Common exam trap
The trap here is assuming that modifying the component's source code automatically changes the cache key when the component is not rebuilt or its inputs are not changed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add the SQL query as a pipeline parameter and pass it to the component as an input.
Vertex AI Pipelines caching uses a fingerprint derived from the component's code and input parameters. When a component's internal logic changes but its inputs and code hash remain the same, the cache is considered valid. To force re-execution without disabling caching entirely, the engineer should expose the changing element—the SQL query—as an input parameter. This changes the cache key only when the query changes, preserving caching for other runs. Rebuilding the image or disabling caching are less precise and more disruptive.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Add the SQL query as a pipeline parameter and pass it to the component as an input.
Why this is correct
Vertex AI Pipelines computes the cache key from the component's inputs (including code and parameters). By promoting the SQL query from hardcoded code to an input parameter, any change to the query alters the input, invalidating the cache and triggering re-execution. This maintains caching benefits for unchanged queries and aligns with pipeline best practices of parameterizing variable logic. It directly solves the problem without globally disabling caching.
- ✗
Set the component's `enable_caching` argument to False in the pipeline definition.
Why it's wrong here
Setting enable_caching to False for that component disables caching for that step, forcing re-execution. However, this permanently disables caching for the component, which is not desired because subsequent unchanged runs would not benefit from cached results. The scenario asks to re-execute with the updated query without disabling caching for the entire pipeline; this approach disables caching for the component, not the entire pipeline, but it also prevents future cache reuse even when the query is unchanged, which is not optimal. A better solution is to surface the query as a parameter.
- ✗
Modify the component's container image tag to a new version and rebuild the pipeline.
Why it's wrong here
Changing the container image tag changes the component's code hash, which would invalidate the cache. However, the scenario states the engineer modified the SQL query inside the component, which is part of the component's code. If the query is embedded in the component's code, rebuilding the image with a new tag would indeed change the cache key. But this is a heavy-handed approach: it requires rebuilding and pushing a new image, and it does not isolate the change to the query parameter. The more precise and maintainable solution is to parameterize the query, not rebuild the entire image.
- ✗
Clear the pipeline's cache by deleting the pipeline run's metadata from Vertex ML Metadata.
Why it's wrong here
Vertex ML Metadata stores pipeline execution metadata and artifacts, but deleting metadata does not directly invalidate the pipeline cache. The cache is based on a fingerprint of the component's inputs and code. Removing metadata could break lineage tracking and is not the supported way to force re-execution. This action is risky and could cause other pipeline runs to lose their lineage. The correct method is to change the component's inputs to alter the cache key.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.