Courseiva

CCNA Pmle Ml Pipelines Questions

6 of 81 questions · Page 2/2 · Pmle Ml Pipelines topic · Answers revealed

76
MCQeasy

A data scientist wants to define a lightweight Python function component in Vertex AI Pipelines using Kubeflow Pipelines SDK v2. Which decorator should be applied to the function to make it a pipeline component?

A.@dsl.pipeline
B.@kfp.v2.components.func_to_component
C.@dsl.component
D.@kfp.dsl.component
AnswerC

The @dsl.component decorator converts a plain Python function into a lightweight KFP v2 component, with the function's signature defining its inputs and outputs. This is the specific mechanism the stem requests for authoring a Python function component in Kubeflow Pipelines SDK v2.

Why this answer

In Kubeflow Pipelines SDK v2, the @dsl.component decorator converts a lightweight Python function into a pipeline component. This is the standard v2 approach for defining function-based components. @dsl.pipeline is used to define the pipeline itself, not individual components.

Exam trap

PMLE often tests the distinction between @dsl.component (defines a component) and @dsl.pipeline (defines the pipeline), and includes legacy v1 decorators like func_to_component to confuse candidates about SDK v2 syntax.

How to eliminate wrong answers

Option A is wrong because @dsl.pipeline decorates the pipeline function that orchestrates components, not the component function itself. Option B is wrong because @kfp.v2.components.func_to_component is the older v1-style API for converting functions to components; in SDK v2 the canonical decorator is @dsl.component. Option D is wrong because @kfp.dsl.component is not the correct v2 decorator syntax; the v2 SDK uses @dsl.component from the kfp.dsl module.

77
MCQmedium

A company uses Vertex AI Pipelines to train models on a daily schedule. The pipeline includes a component that runs a BigQuery query to extract features. The team wants to ensure that if the BigQuery component fails due to transient network errors, the pipeline automatically retries it. How can they configure retries in Vertex AI Pipelines?

A.Deploy the component as a Cloud Function and configure Cloud Functions retry.
B.Wrap the component in a `dsl.If` conditional that checks for failure and re-submits the component.
C.Use Cloud Composer with a task retry policy in Airflow.
D.Set the `retry` parameter of the component to a positive integer, for example `retry=3`.
AnswerD

The retry parameter on a task or component specifies how many times Vertex AI Pipelines re-executes it after failure, so retry=3 automatically reruns the BigQuery component on transient network errors. This satisfies the stem's automatic-retry requirement.

Why this answer

Vertex AI Pipelines natively supports a `retry` parameter on pipeline components. Setting `retry=3` instructs the pipeline to automatically retry the component up to three times if it fails due to transient errors, such as network timeouts. This is the simplest and most direct way to handle retries within the Vertex AI Pipelines orchestration framework.

Exam trap

The trap here is that candidates may confuse Vertex AI Pipelines' native `retry` parameter with external retry mechanisms (Cloud Functions, Airflow) or misuse pipeline control flow constructs like `dsl.If` for retry logic, when the correct approach is a simple parameter on the component definition.

How to eliminate wrong answers

Option A is wrong because deploying the component as a Cloud Function and configuring Cloud Functions retry would move the execution outside of Vertex AI Pipelines, breaking the pipeline's orchestration and monitoring. Option B is wrong because `dsl.If` conditionals are used for conditional execution of components, not for retrying a failed component; they cannot re-submit a component that has already failed. Option C is wrong because Cloud Composer with Airflow is a separate orchestration service that would require migrating the entire pipeline out of Vertex AI Pipelines, adding unnecessary complexity and cost.

78
Multi-Selecthard

A pipeline includes a component that produces a model artifact. The team wants to automatically detect skew between the training data distribution and the serving data distribution. Which three best practices should they implement? (Choose three.)

Select 3 answers
A.Compare statistics using a dedicated component and alert on threshold exceedance
B.Use in-memory data passing for efficiency
C.Compute serving data statistics using a component
D.Disable caching to ensure fresh statistics
E.Pass training data statistics as a Dataset artifact
AnswersA, C, E

A dedicated comparison component computes the statistical distance between training and serving distributions and raises alerts when thresholds are exceeded, satisfying the automatic detection requirement. Isolating this logic keeps skew monitoring reproducible and lets thresholds be tuned without retraining the model.

Why this answer

Option A is correct because skew detection requires a dedicated comparison component that evaluates training versus serving statistics and raises an alert when a configured threshold (e.g., a divergence metric like KL divergence or PSI) is exceeded. Option C is correct because you must compute statistics on the serving data itself, typically via a component that ingests the live inference data and emits a statistics artifact for comparison. Option E is correct because the training data statistics should be persisted and passed as a Dataset artifact so the comparison component can consume them reproducibly across pipeline runs.

Option B is wrong because in-memory data passing does not persist artifacts and is unsuitable for cross-component, cross-run statistical comparison. Option D is wrong because disabling caching does not improve statistical correctness and only adds unnecessary recomputation cost.

79
MCQhard

An ML pipeline runs on Vertex AI and includes a component that uses a third-party library not available in the default Python environment. The team wants to avoid building a custom container image. Which approach should they use?

A.Install the library using pip in the pipeline definition
B.Use a container component with a pre-built image
C.Use the packages_to_install parameter in @dsl.component
D.Add the library to the Vertex AI custom training image
AnswerC

Using `packages_to_install` in `@dsl.component` lets the Vertex AI Pipelines compiler install the third-party library into the component's runtime environment at execution, satisfying the constraint of avoiding a custom container image. The dependency is resolved by pip during task startup rather than being baked into the image.

Why this answer

The `packages_to_install` parameter in the `@dsl.component` decorator allows you to specify a list of third-party Python packages (e.g., via pip) that will be installed at runtime in the component's execution environment, without needing to build a custom container image. This is the recommended approach in Vertex AI Pipelines when you need to use a library not present in the default Python environment, as it avoids the overhead of custom container creation while ensuring the dependency is available for that specific component.

Exam trap

The trap here is that candidates often confuse the `packages_to_install` parameter in Vertex AI's `@dsl.component` with a generic pip install in the pipeline definition, or they assume a pre-built container image avoids custom image building—but in Vertex AI, any container image that includes the library must be custom-built or selected from a registry, which still involves image management overhead. The `packages_to_install` parameter is the native Vertex AI way to install packages without custom containers.

How to eliminate wrong answers

Option A is wrong because `pip install` in the pipeline definition (e.g., in a Python function or YAML) is not a supported mechanism in Vertex AI Pipelines; the pipeline definition itself does not execute shell commands, and dependencies must be declared via the component decorator. Option B is wrong because using a container component with a pre-built image still requires building a custom container image (even if it's pre-built, you must create or select one that includes the library), which contradicts the requirement to avoid building a custom container image. Option D is wrong because adding the library to the Vertex AI custom training image involves creating a custom container image for training, which is a separate process from pipeline components and also requires building a custom image, violating the constraint.

80
MCQhard

A team is building a CI/CD pipeline for an ML model. They want to automatically trigger a Vertex AI pipeline for retraining whenever new training data arrives in a Cloud Storage bucket, but only if a specific Pub/Sub notification is published by a data ingestion process. Which approach meets these requirements with minimal operational overhead?

A.Use Cloud Scheduler to run a job every hour that checks for new files in Cloud Storage and starts the pipeline if new files exist.
B.Configure a Cloud Build trigger that listens to the Pub/Sub topic and executes a build step that submits the pipeline run.
C.Use Eventarc to route the Pub/Sub notification to a Cloud Function that calls the Vertex AI pipeline creation API.
D.Create a Dataflow streaming pipeline that reads from Pub/Sub and triggers the Vertex AI pipeline via a custom sink.
AnswerC

Eventarc natively consumes Pub/Sub messages and delivers them to a Cloud Function, which then invokes the Vertex AI pipeline creation API. This event-driven path avoids polling or custom infrastructure, meeting the trigger condition with minimal operational overhead.

Why this answer

Eventarc can directly listen to a Pub/Sub topic and route matching messages to a Cloud Function, which then calls the Vertex AI pipeline creation API. This serverless approach triggers the pipeline only when the specific Pub/Sub notification is published, meeting the requirement with zero infrastructure to manage and no polling overhead.

Exam trap

The trap here is that candidates may over-engineer the solution by choosing Dataflow (Option D) because it sounds 'streaming' and 'real-time', but the simplest serverless event-driven approach (Eventarc + Cloud Function) meets the requirement with minimal operational overhead.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler polling every hour introduces latency (up to 1 hour) and does not respond to the Pub/Sub notification; it also requires managing a scheduled job and checking for new files, which adds operational overhead and may miss the specific trigger condition. Option B is wrong because Cloud Build triggers are designed for source code changes (e.g., Git commits) and cannot directly listen to a Pub/Sub topic for arbitrary messages; even if configured with a Pub/Sub trigger, Cloud Build is intended for building containers, not for orchestrating ML pipeline runs, and would require extra steps to invoke Vertex AI. Option D is wrong because a Dataflow streaming pipeline is overkill for this simple event-driven trigger; it introduces a persistent streaming job with associated cost and complexity, whereas a lightweight Cloud Function is sufficient and more cost-effective.

81
MCQeasy

A machine learning engineer wants to define a lightweight pipeline component that runs custom Python code without building a container image. Which KFP SDK feature should they use?

A.Importer component
B.Python function component with @dsl.component
C.Container component
D.Vertex AI Training job
AnswerB

The @dsl.component decorator converts a plain Python function into a lightweight pipeline component, letting KFP build the container automatically. This satisfies the stem's requirement to run custom Python code without manually building a container image.

Why this answer

The `@dsl.component` decorator in KFP SDK allows you to define a lightweight Python function component that runs custom code without requiring a container image. It automatically generates a container specification from the function's dependencies, making it ideal for simple, non-containerized pipeline steps.

Exam trap

The trap here is that candidates may confuse 'lightweight' with 'no container at all,' but KFP always runs components in containers; the `@dsl.component` feature automates container creation, not eliminates it.

How to eliminate wrong answers

Option A is wrong because the Importer component is used to import existing artifacts (like datasets or models) into a pipeline, not to run custom Python code. Option C is wrong because a Container component requires you to specify a pre-built container image, which contradicts the requirement of not building a container image. Option D is wrong because Vertex AI Training job is a managed service for running training jobs on Vertex AI, not a lightweight KFP SDK feature for running custom code without containers.

← PreviousPage 2 of 2 · 81 questions total

Ready to test yourself?

Try a timed practice session using only Pmle Ml Pipelines questions.