Courseiva
Monitoring ML Solutions →mediumMultiple Choice

PMLE Monitoring ML Solutions Practice Question

A company wants to track the cost of their Vertex AI prediction endpoint. They use a custom machine type with 1 n1-standard-4 (4 vCPU, 15 GB memory) and 1 NVIDIA T4 GPU. The endpoint is configured for automatic scaling with min=1, max=5 replicas. Which cost monitoring approach should they use?

⚠ Common exam trap

The trap is choosing manual calculation or unrelated tools like Vertex AI Experiments, when the correct approach is to leverage native GCP cost management tools (Cloud Billing + BigQuery export) for accurate and scalable cost monitoring.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Cloud Billing budget alerts and export cost data to BigQuery for analysis.

For tracking Vertex AI prediction endpoint costs, the most comprehensive approach is to use Cloud Billing budget alerts and export cost data to BigQuery for detailed analysis. This allows the company to monitor actual spend, set alerts, and analyze costs by labels, projects, or services, including Vertex AI endpoints with custom machine types and GPUs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use Cloud Billing budget alerts and export cost data to BigQuery for analysis.

    Why this is correct

    Cloud Billing budget alerts notify on spend thresholds, and exporting cost data to BigQuery enables granular analysis by label, including the custom machine type and GPU replicas. This satisfies the stem's requirement to track prediction endpoint costs with autoscaling replicas.

  • ✗

    Calculate cost manually based on replica count and GPU hours from endpoint logs.

    Why it's wrong here

    Manual calculation from logs cannot capture Vertex AI's per-second billing, custom machine pricing, or GPU surcharges, and replica counts fluctuate under autoscaling. It would suit a fixed, single-replica deployment with published flat rates, not a scaling endpoint billed through Cloud Billing.

  • ✗

    Use Vertex AI Experiments to track cost.

    Why it's wrong here

    Vertex AI Experiments records training runs, metrics, and parameters for model evaluation; it holds no billing data and cannot attribute spend to endpoint replicas. It would suit comparing model versions or hyperparameter trials, not monitoring prediction endpoint costs.

  • ✗

    Monitor only the CPU utilisation metrics to infer cost.

    Why it's wrong here

    CPU utilisation indicates load but not spend; GPU hours, per-second custom machine rates, and replica count drive the actual bill. This approach would suit capacity tuning or autoscaling threshold decisions, not cost tracking, which requires Cloud Billing export data.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.