mediumMultiple Select
PDE Practice Question: A data engineering team is building a CI/CD…
A data engineering team is building a CI/CD pipeline for machine learning models using Cloud Build and AI Platform. Which TWO practices are essential for ensuring reproducible and safe model deployments?
⚠ Common exam trap
A common trap in this question is confusing best practices for consistency (such as using the same environment for training and serving) with essential practices for reproducibility and safety (such as version tagging and staged testing) in Google Cloud's AI Platform CI/CD pipelines. Candidates often select Option D because it is a good practice, but it is not explicitly required for reproducibility and safety as defined by the question.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tag each model version with the Git commit hash of the training code.
Option B is correct because tagging each model version with the Git commit hash of the training code creates an immutable link between the deployed artifact and the exact source code that produced it, which is fundamental for reproducibility and auditability in ML CI/CD. Option C is correct because running integration tests against the model on a staging endpoint validates the model's behavior, input/output schema, and serving performance in an environment that mirrors production before promotion, catching regressions safely. Option A is not essential here since Cloud Functions triggering retraining addresses automation of retraining, not reproducibility or safe deployment of models. Option D is not essential and can even be risky, because sharing the same environment for training and serving is not required for reproducibility and may introduce dependency conflicts; custom containers are a separate concern. Option E is incorrect because deploying directly from the development environment with gcloud commands bypasses CI/CD controls, testing, and versioning, which undermines safe and reproducible deployments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud Functions to trigger retraining on new data arrival.
Why it's wrong here
Cloud Functions triggering retraining addresses pipeline automation on data arrival, not reproducibility or deployment safety. It is the right pattern for event-driven orchestration, but the question asks for practices such as versioned artefacts and staged validation before promotion.
- ✓
Tag each model version with the Git commit hash of the training code.
Why this is correct
Tagging each model version with the training code's Git commit hash creates a traceable link between artefact and source, so any deployment can be reproduced exactly. This satisfies the reproducibility requirement of the CI/CD pipeline by pinning the code revision that produced the model.
- ✓
Run integration tests against the model on a staging endpoint before promoting to production.
Why this is correct
Running integration tests against a staging endpoint exercises the deployed model's real serving path, including container, dependencies and I/O, before production promotion. This satisfies the safe-deployment requirement by catching serving failures that unit tests on training code cannot detect.
- ✗
Use the same environment for training and serving, possibly via custom containers.
Why it's wrong here
Sharing one environment for training and serving couples the two lifecycles, so serving changes can break training reproducibility and vice versa. Identical custom containers are correct for guaranteeing inference parity, but the pipeline still needs separate, versioned training and serving environments.
- ✗
Directly deploy from the development environment using gcloud commands.
Why it's wrong here
Deploying straight from development with gcloud commands bypasses the Cloud Build pipeline, so no versioned artefact, testing or approval gate exists. Direct CLI deployment suits rapid prototyping in a sandbox, not reproducible, safe promotion to production.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.