Courseiva

CCNA Pmle Low Code Ml Questions

68 questions · Pmle Low Code Ml topic · All types, answers revealed

1
MCQmedium

A company needs to extract key fields from scanned invoices, such as invoice number and total amount, with high accuracy. They want a managed service and plan to use human review for low-confidence results. Which combination of services should they use?

A.Vision API and Natural Language API
B.Document AI and Human-in-the-Loop
C.Translation API and AutoML Vision
D.BigQuery ML and Vertex AI Prediction
AnswerB

Document AI extracts invoice fields via prebuilt or custom models, returning per-field confidence scores. Human-in-the-Loop routes only low-confidence extractions to reviewers, satisfying the stem's accuracy requirement while keeping the service fully managed. This pairing directly matches the stated need for human review of uncertain results.

Why this answer

Document AI is Google Cloud's managed document understanding service with specialized parsers (Invoice, Expense, Form) that extract structured fields like invoice number and total amount with high accuracy. Human-in-the-Loop (HITL) integrates directly with Document AI to route low-confidence predictions to human reviewers, whose corrections feed back to improve the processor. Together they satisfy the managed-service and human-review requirements in one pipeline.

Exam trap

PMLE often tests the confusion between generic Vision/NLP APIs and purpose-built Document AI — candidates assume any OCR service can extract invoice fields, missing that schema-aware extraction plus a native human-review loop is the differentiator.

How to eliminate wrong answers

Option A is wrong because Vision API performs generic OCR/image labeling and Natural Language API does entity/sentiment analysis — neither understands invoice schema or returns structured invoice fields. Option C is wrong because Translation API only converts languages and AutoML Vision is a custom image classifier, not a document field extractor; you would have to build and label everything yourself. Option D is wrong because BigQuery ML and Vertex AI Prediction are general ML training/serving tools — they do not provide pre-built document parsing or a human-review workflow.

2
Multi-Selectmedium

A company wants to analyze videos to detect objects and track their movement over time. Which TWO Google Cloud services are suitable for this task?

Select 2 answers
A.AutoML Vision
B.Speech-to-Text
C.AutoML Video
D.Video Intelligence API
E.Natural Language API
AnswersC, D

AutoML Video lets you train a custom model on labelled video to detect and track objects specific to your domain, then serve predictions on new footage. This satisfies the movement-tracking requirement when pre-trained labels are insufficient.

Why this answer

AutoML Video supports object tracking, and Video Intelligence API provides both object detection and tracking. AutoML Vision is for images only, Natural Language for text, and Speech-to-Text for audio.

3
MCQmedium

A data engineer wants to use BigQuery ML to train a model that predicts customer churn using a table with customer features and a label column. They want to use a deep neural network. Which model type should they specify?

A.BOOSTED_TREE_CLASSIFIER
B.LOGISTIC_REG
C.DNN_CLASSIFIER
D.DNN_REGRESSOR
AnswerC

DNN_CLASSIFIER fits because churn prediction is binary classification, and the stem specifies a deep neural network. BigQuery ML's DNN_CLASSIFIER trains a feed-forward neural network with a logistic output layer, satisfying both the label-column requirement and the requested architecture, unlike DNN_REGRESSOR, which predicts continuous values.

Why this answer

The DNN_CLASSIFIER model type in BigQuery ML is specifically designed for classification tasks using a deep neural network architecture. Since the problem is predicting customer churn (a binary classification problem) and the data engineer explicitly wants to use a deep neural network, DNN_CLASSIFIER is the appropriate choice.

Exam trap

The trap is that candidates may confuse DNN_REGRESSOR with DNN_CLASSIFIER, but BigQuery ML uses different model types for regression vs. classification. DNN_CLASSIFIER is used for classification tasks like churn prediction, while DNN_REGRESSOR is for continuous values.

How to eliminate wrong answers

Option A is wrong because BOOSTED_TREE_CLASSIFIER uses gradient-boosted decision trees, not a deep neural network, so it does not meet the requirement for a DNN model. Option B is wrong because LOGISTIC_REG is a logistic regression model, which is a linear classifier and not a deep neural network. Option D is wrong because DNN_REGRESSOR is used for regression tasks (predicting continuous values), not for classification tasks like churn prediction.

4
MCQmedium

A healthcare company wants to build a model to predict patient readmission risk using structured data in BigQuery. They have a dataset with 100,000 rows and 30 features, including numerical and categorical variables. They require a model that provides explainable predictions and can be trained quickly. They decide to use BigQuery ML. Which model type should they choose?

A.Matrix factorization
B.Logistic regression
C.K-means clustering
D.Deep neural network (DNN)
AnswerB

Logistic regression is a linear model for binary classification that provides interpretable coefficients, making it suitable for explainable predictions. It trains quickly on structured data and handles both numerical and categorical features after preprocessing. For predicting patient readmission (a binary outcome), logistic regression is a strong choice in BigQuery ML, especially when explainability is required. It also supports regularization to prevent overfitting.

Why this answer

Logistic regression is the best choice because it is interpretable, fast to train, and effective for binary classification with structured data. It provides coefficients that indicate feature importance, which helps explain predictions. BigQuery ML supports logistic regression with options for regularization, and it can handle categorical variables automatically.

This aligns with the requirement for explainable predictions and quick training.

Exam trap

The trap here is assuming that a more complex model like a deep neural network is always better, ignoring the need for explainability.

5
MCQeasy

A data analyst wants to build a binary classification model to predict customer churn using SQL queries in BigQuery. Which BigQuery ML model type should they use?

A.MATRIX_FACTORIZATION
B.LINEAR_REG
C.LOGISTIC_REG
D.K_MEANS
AnswerC

LOGISTIC_REG performs binary logistic regression inside BigQuery, outputting probabilities for two-class labels such as churn or no churn. It satisfies the stem's requirement for a binary classification model built directly through SQL, with no data export or external tooling needed.

Why this answer

LOGISTIC_REG is the BigQuery ML model type for binary classification, predicting a binary outcome such as churn (yes/no). It uses logistic regression to estimate the probability of the binary outcome. This is the correct choice for predicting customer churn.

Exam trap

PMLE often tests the confusion between regression and classification model types, leading candidates to pick LINEAR_REG for binary outcomes or K_MEANS for supervised tasks.

How to eliminate wrong answers

Option A is wrong because MATRIX_FACTORIZATION is used for recommendation systems and collaborative filtering, not binary classification. Option B is wrong because LINEAR_REG is for linear regression, predicting continuous numeric values, not binary outcomes. Option D is wrong because K_MEANS is for clustering, an unsupervised learning technique, not classification.

6
MCQmedium

A company wants to analyze customer reviews for sentiment (positive, negative, neutral) using a pre-trained model with no training. They have text data stored in BigQuery. Which Google Cloud service should they use?

A.Translation API
B.Speech-to-Text API
C.Natural Language API
D.AutoML NLP
AnswerC

The Natural Language API provides pre-trained sentiment analysis, returning positive, negative and neutral scores without any custom training. It integrates directly with BigQuery data, satisfying the requirement for a pre-trained model applied to stored review text.

Why this answer

The Natural Language API is a pre-trained service that provides sentiment analysis out of the box, requiring no training. It can directly process text data from BigQuery (via integration or export) to classify sentiment as positive, negative, or neutral. This matches the requirement of using a pre-trained model with no training.

Exam trap

PMLE often tests the distinction between pre-trained APIs and custom training services; candidates may mistakenly choose AutoML NLP thinking it's needed for sentiment, but AutoML requires training data.

How to eliminate wrong answers

Option A is wrong because the Translation API translates text between languages, not analyze sentiment. Option B is wrong because Speech-to-Text transcribes audio, not text sentiment. Option D is wrong because AutoML NLP requires custom training with labeled data, which contradicts the 'no training' requirement.

7
Multi-Selecteasy

A company wants to transcribe audio from customer service calls and then analyze the sentiment of the transcribed text. Which TWO Google Cloud services should they use?

Select 2 answers
A.Natural Language API
B.Document AI
C.Speech-to-Text
D.Translation API
E.Vision API
AnswersA, C

The Natural Language API performs sentiment analysis on text, satisfying the requirement to analyse the transcribed call content. Speech-to-Text handles the audio transcription; Natural Language API then classifies each transcript's sentiment. Together they cover both stages of the pipeline, making this one of the two required services.

Why this answer

Speech-to-Text (option C) is correct because it converts audio from the customer service calls into written text, which is the required transcription step. Natural Language API (option A) is correct because it performs sentiment analysis on text, allowing the company to analyze the sentiment of the transcribed call content. Document AI (option B) is not appropriate here because it processes documents and forms rather than audio or general sentiment analysis.

Translation API (option D) only translates text between languages and does not transcribe audio or analyze sentiment. Vision API (option E) analyzes images, so it cannot handle audio transcription or text sentiment analysis.

Exam trap

PMLE often tests the combination of services for multi-step tasks; candidates might choose Document AI for transcription or Translation API for sentiment, but these are incorrect service mappings.

8
Multi-Selectmedium

A marketing team wants to build a model to predict which customers are likely to churn. They have a BigQuery table with customer demographics, usage metrics, and a binary churn label. They want to use BigQuery ML and need to evaluate the model's performance. Which two statements are true regarding model evaluation in BigQuery ML? (Choose two.)

Select 2 answers
A.ML.FEATURE_INFO returns the importance of each feature in the model.
B.ML.PREDICT automatically calculates the model's accuracy on the input data.
C.ML.TRAINING_INFO provides the evaluation metrics for the trained model.
D.ML.CONFUSION_MATRIX can be used to visualize the confusion matrix of a classification model.
E.ML.EVALUATE returns metrics such as precision, recall, accuracy, and AUC for classification models.
AnswersD, E

ML.CONFUSION_MATRIX generates a confusion matrix showing true positives, true negatives, false positives, and false negatives. This is useful for understanding the types of errors the model makes, which is critical when the cost of false positives and false negatives differs, as in churn prediction.

Why this answer

ML.EVALUATE computes standard classification metrics, and ML.CONFUSION_MATRIX provides a detailed breakdown of prediction outcomes. Both are essential for assessing a churn model. The other functions serve different purposes: ML.PREDICT for scoring, ML.TRAINING_INFO for training details, and ML.FEATURE_INFO for feature statistics.

Exam trap

The trap here is confusing prediction with evaluation; ML.PREDICT does not evaluate, and ML.TRAINING_INFO does not provide classification metrics.

9
MCQeasy

A small marketing team has a CSV file of 2,000 labeled customer support tickets (each with a category such as 'billing' or 'technical'). They have no ML engineers and want a fully managed, low-code way to train a text classification model that they can later call from their internal web app. Which Google Cloud service should they use?

A.BigQuery ML with a CREATE MODEL statement using the LOGISTIC_REG model type
B.Vertex AI Pipelines with a custom Kubeflow component for text preprocessing
C.Cloud Natural Language API classifyText method
D.Vertex AI AutoML text classification
AnswerD

AutoML text classification is a managed, low-code service that trains on a labeled CSV in a Vertex AI dataset and exposes a prediction endpoint for app integration. It handles model selection and tuning, matching the team's lack of ML engineers and need for a callable API.

Why this answer

AutoML text classification is purpose-built for teams with labeled text and no ML expertise: it ingests a CSV, trains, tunes, and serves a model behind a managed endpoint. The other services either require custom code, expect structured numeric features, or rely on a fixed taxonomy that cannot represent the team's custom ticket categories.

Exam trap

The trap here is assuming a general NLP API with a classify method can be trained on custom labels, when it actually applies only a predefined taxonomy.

10
Multi-Selecthard

A retail company wants to implement a recommendation system using Recommendations AI. They need to generate personalized recommendations for users based on their browsing history and purchase behavior. Which THREE recommendation types are available in Recommendations AI?

Select 3 answers
A.trending-now
B.recommended-for-you
C.others-you-may-like
D.frequently-bought-together
E.most-popular
AnswersB, C, D

Recommended-for-you is a genuine Recommendations AI type that personalises suggestions per user from their browsing history and past behaviour. This directly matches the stem's requirement to generate personalised recommendations based on browsing and purchase activity.

Why this answer

Recommendations AI offers several built-in recommendation types, and recommended-for-you (B) is correct because it produces personalized product suggestions for a user based on that user's own browsing history, purchase behavior, and other interaction events. others-you-may-like (C) is also correct because it generates personalized recommendations related to a specific item the user is currently viewing, which fits the retail scenario of tailoring suggestions from user behavior. frequently-bought-together (D) is correct because it recommends complementary products commonly purchased alongside a given item, a standard Recommendations AI type for cross-sell use cases. The unmarked options do not belong: trending-now (A) and most-popular (E) are not valid Recommendations AI recommendation type identifiers, as the platform's types are named like recommended-for-you, others-you-may-like, frequently-bought-together, and similar-item, rather than those generic popularity labels.

Exam trap

PMLE often tests the distinction between valid recommendation types in Recommendations AI and generic recommendation strategies like 'trending-now' or 'most-popular', which are not available as predefined types.

11
MCQhard

A company has an existing TensorFlow model for fraud detection that they want to use for predictions in BigQuery. They want to call the model from SQL queries without moving data out of BigQuery. How should they deploy the model?

A.Import the TensorFlow model directly into BigQuery ML
B.Deploy the model to Vertex AI Prediction and use a remote model in BigQuery ML
C.Export BigQuery data to Cloud Storage and use AI Platform Prediction
D.Use AutoML Tables to retrain the model in BigQuery ML
AnswerA

Why this answer

BigQuery ML (BQML) natively supports importing TensorFlow models directly, allowing you to use them for predictions via SQL without moving data out of BigQuery. This is the simplest and most efficient approach because it eliminates the need for external services or data export, leveraging BQML's built-in `CREATE MODEL` statement with the `OPTIONS(model_type = 'TENSORFLOW')` clause.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing Vertex AI or AI Platform, not realizing that BigQuery ML has native TensorFlow support, which is the most direct and low-code way to meet the requirement of keeping data in BigQuery.

How to eliminate wrong answers

Option B is wrong because deploying to Vertex AI Prediction and using a remote model adds unnecessary complexity and latency; while it works, it requires setting up a remote model and a connection, which is not the simplest or most direct method when BQML directly supports TensorFlow imports. Option C is wrong because exporting BigQuery data to Cloud Storage and using AI Platform Prediction moves data out of BigQuery, violating the requirement to keep data in BigQuery and adding extra steps and cost. Option D is wrong because AutoML Tables retrains the model from scratch, which does not reuse the existing TensorFlow model and may produce different results, whereas the requirement is to use the existing model as-is.

12
MCQmedium

A logistics company wants to classify shipping documents into categories such as invoice, packing slip, and bill of lading. They have a small set of labeled documents (about 50 per category) and want to use a low-code approach. They need a model that can be trained quickly and deployed for online predictions. Which Google Cloud service should they use?

A.AutoML Natural Language
B.Document AI
C.Natural Language API
D.BigQuery ML
AnswerA

AutoML Natural Language allows training custom text classification models with as few as 50 labeled documents per category. It provides a low-code UI and API, and can deploy models for online predictions. It is designed for exactly this scenario: small labeled datasets and quick training.

Why this answer

AutoML Natural Language is designed for custom text classification with small labeled datasets. It provides a low-code environment, automatically handles preprocessing, and allows easy deployment for online predictions, making it ideal for classifying shipping documents.

Exam trap

The trap here is assuming that the pre-trained Natural Language API can be customized, but it only offers general categories, not user-defined ones.

13
MCQmedium

An organisation wants to use Document AI to process contracts but requires human review for high-risk clauses. Which feature should they enable?

A.Human-in-the-Loop (HITL)
B.Batch Processing
C.Online Prediction
D.AutoML Training
AnswerA

Human-in-the-Loop routes low-confidence or high-risk clause extractions to human reviewers before finalising results, satisfying the stem's requirement for mandatory human review of high-risk clauses. Document AI's confidence thresholds trigger this escalation automatically, so contracts proceed without manual intervention except where risk demands it.

Why this answer

Human-in-the-Loop (HITL) is the correct feature because it allows Document AI to automatically process contracts while routing high-risk clauses to human reviewers for validation. This balances automation efficiency with the need for expert oversight on sensitive content, which is a core requirement for compliance-driven document processing.

Exam trap

The trap here is that candidates confuse HITL with AutoML Training, thinking that training a model with human-labeled data is the same as having a human review live predictions, but HITL is a runtime workflow, not a training process.

How to eliminate wrong answers

Option B (Batch Processing) is wrong because it handles large volumes of documents asynchronously but does not include any mechanism for human review or intervention on specific clauses. Option C (Online Prediction) is wrong because it provides real-time predictions on individual documents but lacks the built-in workflow to pause and escalate high-risk clauses to a human. Option D (AutoML Training) is wrong because it is used to train custom models on labeled data, not to manage human review workflows during inference.

14
Multi-Selecthard

A media company wants a low-code pipeline that ingests uploaded video files, detects scenes and on-screen text, and stores structured metadata for search. They prefer managed services and minimal custom code. Which TWO Google Cloud capabilities should they combine? (Choose two.)

Select 2 answers
A.Cloud Vision API for detecting labels and text in sampled video frames
B.Vertex AI Vision for building a custom object-tracking pipeline with trained detectors
C.Cloud Data Fusion for building a visual ETL pipeline for video ingestion
D.Cloud Storage and Pub/Sub to trigger processing when new videos arrive
E.Video Intelligence API for shot change detection and text detection
AnswersD, E

A Cloud Storage bucket receiving uploads can publish object notifications to Pub/Sub, which then triggers the video analysis step. This serverless eventing is the standard low-code way to automate ingestion so each new video is processed and its metadata stored without manual intervention.

Why this answer

The Video Intelligence API supplies managed shot change and text detection on video, producing the scene and on-screen text metadata needed for search. Pairing it with Cloud Storage object notifications to Pub/Sub creates an event-driven trigger so each upload is analyzed automatically. Together they deliver a managed, low-code pipeline without custom frame extraction or model training.

Exam trap

The trap here is substituting frame-by-frame Cloud Vision API calls for the video-native Video Intelligence API, which ignores the extra code required to sample frames and align timestamps.

15
MCQhard

A team trained a TensorFlow model locally and wants to deploy it to BigQuery ML for predictions without retraining. They have exported the SavedModel to Cloud Storage. Which statement is correct?

A.They need to convert the model to a BigQuery ML native format first.
B.They can create a model using CREATE MODEL with model_type='tensorflow' and the path to the SavedModel.
C.They must first retrain the model using ML.TRAIN on BigQuery.
D.They can use ML.PREDICT directly on the SavedModel in Cloud Storage.
AnswerB

CREATE MODEL with model_type='tensorflow' imports an existing SavedModel directly from Cloud Storage, so BigQuery ML serves predictions without retraining. This satisfies the stem's constraint: the locally trained TensorFlow artefact is deployed as-is, with no data or training step required in BigQuery.

Why this answer

BigQuery ML supports importing TensorFlow models in SavedModel format directly using CREATE MODEL with model_type='tensorflow' and the path to the SavedModel in Cloud Storage. This allows you to use the model for prediction in BigQuery without retraining or converting to a native format. The model is imported as a BigQuery ML model and can be used with ML.PREDICT.

Exam trap

The trap is assuming that BigQuery ML requires models to be trained within BigQuery or converted to a proprietary format. In fact, it supports direct import of TensorFlow SavedModels.

How to eliminate wrong answers

Option A is wrong because no conversion to a native BigQuery ML format is needed; TensorFlow SavedModels are supported directly. Option C is wrong because retraining is not required; the whole point is to use the pre-trained model. Option D is wrong because ML.PREDICT cannot be used directly on a SavedModel in Cloud Storage; you must first create a BigQuery ML model that references the SavedModel.

16
MCQmedium

A data analyst wants to train a binary classification model in BigQuery ML on a dataset of 10 million rows with 50 features. They need to evaluate the model's performance on a held-out test set. Which sequence of SQL statements should they run?

A.CREATE MODEL then ML.FEATURE_IMPORTANCE
B.ML.TRAIN then ML.EVALUATE
C.CREATE MODEL then ML.PREDICT
D.CREATE MODEL then ML.EVALUATE
AnswerD

BigQuery ML trains the model with CREATE MODEL, which handles the 10-million-row, 50-feature dataset natively in SQL. ML.EVALUATE then scores that trained model against the held-out test set, returning precision, recall, AUC and related metrics, satisfying the evaluation requirement without exporting data.

Why this answer

To train and evaluate a model in BigQuery ML, you use CREATE MODEL to train the model, and then ML.EVALUATE to assess its performance on a held-out dataset. CREATE MODEL automatically splits the data into training and evaluation sets if specified, or you can use a separate table for evaluation. ML.EVALUATE returns metrics like accuracy, precision, recall, etc.

This sequence is standard for model development in BigQuery ML.

Exam trap

The trap is confusing ML.EVALUATE with ML.PREDICT or thinking that ML.TRAIN exists. Candidates might also think they need to use ML.FEATURE_IMPORTANCE for evaluation.

How to eliminate wrong answers

Option A is wrong because ML.FEATURE_IMPORTANCE is used to understand feature importance, not to evaluate model performance. Option B is wrong because ML.TRAIN is not a valid BigQuery ML function; training is done via CREATE MODEL. Option C is wrong because ML.PREDICT is for generating predictions, not for evaluating performance on a test set.

17
MCQhard

A healthcare organization wants to build a model to predict patient readmission risk using structured electronic health record (EHR) data. They need to train a model using SQL in BigQuery, but they also want to leverage AutoML's ability to automatically search for the best architecture. Which approach should they take?

A.Use a pre-built Vision API model via BigQuery ML remote model
B.Use BigQuery ML with the AUTOML_CLASSIFIER model type
C.Use AutoML Tables with Vertex AI and export predictions
D.Use BigQuery ML with a DNN_CLASSIFIER and manual hyperparameter tuning
AnswerB

BigQuery ML's AUTOML_CLASSIFIER model type trains directly on BigQuery data using SQL while running AutoML architecture search and hyperparameter tuning behind the scenes. This satisfies both the SQL-in-BigQuery requirement and the demand for automated model selection on structured EHR data.

Why this answer

BigQuery ML's AUTOML_CLASSIFIER model type automatically performs architecture search and hyperparameter tuning, making it ideal for users who want to leverage AutoML capabilities directly within SQL on structured EHR data. This approach avoids manual model selection while staying entirely within BigQuery's SQL interface, which is the stated requirement.

Exam trap

The trap here is that candidates confuse AutoML Tables (a separate Vertex AI service) with BigQuery ML's built-in AUTO model type, assuming they must export data to use AutoML, when in fact BigQuery ML provides AutoML capabilities directly within SQL.

How to eliminate wrong answers

Option A is wrong because Vision API is designed for image analysis, not structured EHR data, and BigQuery ML remote models require a pre-built API endpoint, not AutoML architecture search. Option C is wrong because AutoML Tables (now Vertex AI Tabular) is a separate service that requires exporting data out of BigQuery and does not allow training via SQL in BigQuery. Option D is wrong because DNN_CLASSIFIER with manual hyperparameter tuning contradicts the requirement to 'automatically search for the best architecture' — it requires explicit user-specified parameters and does not perform automated architecture search.

18
MCQhard

A media company wants to automatically moderate user-uploaded videos by detecting explicit content (e.g., violence, adult material). They need a solution that integrates with their video processing pipeline and scales to millions of videos. Which approach should they take?

A.Use Video Intelligence API with explicit content detection
B.Use AutoML Video to train a custom explicit content detection model
C.Use Natural Language API on video transcripts
D.Use Vision API to analyze each video frame
AnswerA

The Video Intelligence API provides explicit content detection purpose-built for video, analysing frames and audio for violence and adult material. It integrates into automated pipelines and scales elastically, satisfying the stem's requirement to moderate millions of videos without building custom models.

Why this answer

Video Intelligence API provides explicit content detection specifically designed to identify violence, adult material, and other explicit content in videos. It is a managed service that scales automatically and can be integrated into video processing pipelines via its API. This is the most direct and scalable solution for moderating user-uploaded videos.

Exam trap

The trap is overcomplicating the solution by considering custom model training (AutoML) or using image analysis frame-by-frame. The exam expects you to know that Video Intelligence API has built-in explicit content detection.

How to eliminate wrong answers

Option B is wrong because AutoML Video requires training a custom model, which is time-consuming and unnecessary when a pre-trained model for explicit content detection already exists. Option C is wrong because Natural Language API analyzes text, not video content; it would only work on transcripts and miss visual explicit content. Option D is wrong because Vision API analyzes individual images, not videos; processing each frame would be inefficient and costly, and it lacks temporal context.

19
MCQhard

A company has a TensorFlow model trained outside of Google Cloud and wants to use it for online predictions on Vertex AI. They have saved the model in SavedModel format. What is the most efficient way to deploy this model?

A.Import the model into BigQuery ML using CREATE MODEL with model_type='TENSORFLOW'
B.Use Vertex AI AutoML Tables to retrain the model
C.Use Cloud Functions to run the model for each prediction request
D.Upload the saved model to Vertex AI and create an endpoint for online predictions
AnswerD

Vertex AI accepts SavedModel artefacts directly, so uploading the existing model and deploying it to an endpoint avoids retraining or conversion. This is the most efficient route to online predictions for a TensorFlow model trained outside Google Cloud.

Why this answer

The most efficient way to deploy a TensorFlow SavedModel for online predictions on Vertex AI is to upload the SavedModel to Vertex AI and create an endpoint. Vertex AI supports importing custom models in SavedModel format and deploying them to endpoints for low-latency online predictions. This leverages Vertex AI's managed infrastructure for scaling and serving.

Exam trap

The trap is confusing deployment targets: BigQuery ML is for SQL predictions, Cloud Functions for lightweight tasks, and AutoML for training. The correct choice is Vertex AI for custom model serving.

How to eliminate wrong answers

Option A is wrong because importing into BigQuery ML is for SQL-based predictions, not for online serving; it also may not support all TensorFlow ops. Option B is wrong because AutoML Tables is for training models from tabular data, not for deploying existing TensorFlow models. Option C is wrong because Cloud Functions is not designed for serving ML models at scale and would have cold starts and limitations.

20
MCQmedium

A logistics company wants to classify shipping documents into categories (invoice, packing slip, bill of lading) using a custom model with minimal code. They have labeled training images. Which Google Cloud service is most appropriate?

A.Vertex AI AutoML Tables
B.AutoML Vision for image classification
C.Document AI custom extractor
D.Cloud Vision API with label detection
AnswerB

AutoML Vision trains a custom image classification model from labelled images through a graphical interface, requiring minimal code. It fits the three document categories and the labelled training images, unlike pre-built Vision API which cannot learn bespoke classes.

Why this answer

AutoML Vision for image classification is designed to train custom image classification models with minimal code using labeled images. It automatically handles data preprocessing, model selection, and hyperparameter tuning, making it ideal for classifying shipping documents from images. The other services are either for tabular data, document extraction, or pre-trained label detection without custom training.

Exam trap

The trap is confusing Document AI with AutoML Vision; candidates might think Document AI can classify documents, but it is primarily for extraction, while AutoML Vision is for custom image classification.

How to eliminate wrong answers

Option A is wrong because Vertex AI AutoML Tables is for tabular data, not images. Option C is wrong because Document AI custom extractor is for extracting structured data from documents, not for classifying document types. Option D is wrong because Cloud Vision API with label detection uses pre-trained models to detect general labels, but it cannot be customized to classify specific document categories like invoice, packing slip, or bill of lading.

21
MCQhard

A financial institution needs to extract structured data from scanned PDFs of loan applications, including text fields and tables. They require a human review step for high-risk applications. Which Google Cloud service and configuration should they use?

A.Document AI with a form parser processor and enable Human-in-the-Loop for high-risk applications
B.Document AI with a custom extractor processor and use Cloud Functions for human review
C.Cloud Vision API to detect text and tables, then send to Cloud Dataflow for processing
D.Vertex AI AutoML Vision to train a custom model for document parsing
AnswerA

Document AI's form parser processor extracts key-value pairs and tables from scanned PDFs via OCR and layout-aware parsing, satisfying the structured-data requirement. Enabling Human-in-the-Loop routes high-risk applications to human reviewers, directly meeting the mandated review step for those cases.

Why this answer

Document AI's Form Parser processor is purpose-built to extract structured key-value pairs and tables from forms like loan applications, and it natively integrates with Human-in-the-Loop (HITL) to route low-confidence or high-risk documents to human reviewers. This combination directly satisfies both the extraction and human review requirements without custom code.

Exam trap

PMLE often tests the misconception that any Document AI processor supports HITL — in reality, HITL is a specific feature that must be enabled and is best paired with Form Parser or custom extractors, not with generic OCR or Vision API.

How to eliminate wrong answers

Option B is wrong because while a custom extractor can be trained, it does not natively provide the HITL workflow — using Cloud Functions for review is a manual, non-integrated workaround that lacks Document AI's built-in confidence-based routing. Option C is wrong because Cloud Vision API only performs OCR/text detection and cannot natively parse structured form fields or tables into key-value pairs; Dataflow is a streaming/batch pipeline, not a document parser. Option D is wrong because AutoML Vision is an image classification/object detection tool, not designed for structured document field extraction from PDFs.

22
Multi-Selecthard

A company is building a document processing pipeline using Document AI to extract data from invoices. They want to ensure high accuracy and handle edge cases where the model may be uncertain. Which THREE steps should they include in their pipeline?

Select 3 answers
A.Regularly retrain the processor using human-verified data
B.Use the pre-built invoice parser without any modifications
C.Use AutoML Vision to classify invoice types
D.Enable Human-in-the-Loop (HITL) to review documents with low confidence scores
E.Use a custom processor trained on their specific invoice format
AnswersA, D, E

Regular retraining with human-verified data directly addresses the accuracy constraint by feeding corrected edge-case extractions back into the processor, letting it learn the specific invoice variations it previously misread. This closes the loop on uncertain predictions, since human review resolves ambiguity that the model alone cannot, progressively raising extraction confidence across the pipeline.

Why this answer

Option A is correct because regularly retraining the processor with human-verified data continuously improves extraction accuracy and adapts the model to new invoice variations and edge cases over time. Option D is correct because enabling Human-in-the-Loop (HITL) routes documents with low confidence scores to human reviewers, ensuring uncertain or edge-case extractions are validated and corrected before entering downstream systems. Option E is correct because a custom processor trained on the company's specific invoice format captures their unique layouts, fields, and terminology, yielding higher accuracy than a generic model.

Option B is not appropriate because using the pre-built invoice parser without modification offers no tuning for the company's specific formats and provides no mechanism for handling uncertainty. Option C is not appropriate because AutoML Vision is an image classification service, not a document entity-extraction tool, and classifying invoice types does not extract the required field data or address low-confidence edge cases.

Exam trap

PMLE often tests the misconception that a pre-built parser is sufficient for all use cases — candidates overlook that custom training and HITL are required for high accuracy on domain-specific documents.

23
MCQeasy

A data scientist wants to evaluate the performance of a BigQuery ML classification model on a test dataset. Which function should they use?

A.ML.PREDICT
B.ML.FEATURE_IMPORTANCE
C.ML.EVALUATE
D.ML.TRAIN
AnswerC

ML.EVALUATE computes standard classification metrics such as precision, recall, accuracy, F1 score, log loss and ROC AUC against a labelled dataset. It is the BigQuery ML function designed for assessing an already-trained model's performance, unlike ML.PREDICT, which only returns predictions.

Why this answer

ML.EVALUATE is the correct function because it computes classification metrics (e.g., precision, recall, accuracy, F1 score, ROC AUC) directly on a trained BigQuery ML model using a provided test dataset or evaluation input. This is the dedicated function for assessing model performance after training, aligning with the task of evaluating a classification model on held-out test data.

Exam trap

Google often tests the distinction between prediction (ML.PREDICT) and evaluation (ML.EVALUATE), trapping candidates who confuse generating outputs with measuring performance, especially when the question mentions 'evaluate performance' but the candidate fixates on 'predict' as the primary ML function.

How to eliminate wrong answers

Option A is wrong because ML.PREDICT is used to generate predictions (class labels or probabilities) on new data, not to compute evaluation metrics like accuracy or precision. Option B is wrong because ML.FEATURE_IMPORTANCE is used to retrieve feature weights or importance scores from a trained model (e.g., for interpretability), not to evaluate overall model performance on a test set. Option D is wrong because ML.TRAIN is used to initiate the training process of a BigQuery ML model, not to evaluate an already trained model on test data.

24
MCQeasy

A marketing team wants to automatically categorize customer feedback emails into topics such as 'billing', 'technical support', or 'general inquiry'. They have a dataset of 5,000 labeled emails and want to build a custom model with minimal coding effort. Which Google Cloud service should they use?

A.Dialogflow CX
B.Vision API
C.Natural Language API
D.AutoML Natural Language
AnswerD

AutoML Natural Language allows training custom text classification models with minimal coding, using labeled data. It handles preprocessing, training, and deployment automatically. With 5,000 labeled emails, it can achieve high accuracy and is ideal for teams with limited ML expertise. It integrates with other Google Cloud services and provides a user-friendly interface, making it the best fit for this low-code scenario.

Why this answer

AutoML Natural Language is purpose-built for custom text classification with minimal coding. It allows training on labeled data to recognize specific categories like billing or technical support. The other services either provide only pre-trained general models or are designed for different modalities, making AutoML Natural Language the correct choice for this low-code custom classification task.

Exam trap

The trap here is confusing the pre-trained Natural Language API with AutoML Natural Language; the former cannot be customized for specific topics without training, while the latter is designed for custom classification.

25
Multi-Selectmedium

A retail company wants to build a low-code ML solution to predict customer lifetime value (CLV) using historical transaction data stored in BigQuery. They have limited ML expertise and want to use BigQuery ML. Which two steps are necessary to train and evaluate a model using BigQuery ML? (Choose two.)

Select 2 answers
A.Use ML.PREDICT to generate predictions on new data.
B.Manually split the data into training and test sets using a CREATE TABLE statement.
C.Use ML.EVALUATE to assess the model's performance on a test dataset.
D.Export the model to a TensorFlow SavedModel for deployment.
E.Create a model using the CREATE MODEL statement with the appropriate model type.
AnswersC, E

ML.EVALUATE is a function that computes evaluation metrics for a trained model, such as RMSE for regression. It is essential to assess model performance and ensure it meets business needs. You run it against a dataset not used in training. This step is necessary for model validation and is performed via SQL, maintaining the low-code approach.

Why this answer

Training a model in BigQuery ML requires the CREATE MODEL statement, which defines and trains the model using SQL. Evaluation is performed with ML.EVALUATE to compute metrics like RMSE or accuracy. These two steps are essential for building and assessing a model.

Other steps like exporting or predicting are optional or occur after evaluation, and manual data splitting is unnecessary due to automatic splitting.

Exam trap

The trap here is thinking that data must be manually split or that model export is required for training; BigQuery ML handles splitting automatically and export is only for external deployment.

26
MCQmedium

A retail company wants to forecast daily sales for each of its 500 stores for the next 90 days. They have three years of historical daily sales data stored in BigQuery, including promotions, holidays, and store attributes. The data science team has minimal ML expertise and wants to use SQL to build and deploy the model with minimal coding. Which approach should they use?

A.Build a custom LSTM model in Vertex AI using TensorFlow, training on the historical sales data.
B.Use AutoML Forecasting in Vertex AI, exporting the data from BigQuery to Cloud Storage.
C.Train a linear regression model in BigQuery ML with store ID as a feature and date as a numeric feature.
D.Create an ARIMA_PLUS model in BigQuery ML using the sales time series and specify the store ID as the time series identifier.
AnswerD

ARIMA_PLUS is designed for univariate time series forecasting and can handle multiple time series by specifying a time series ID column. It automatically handles seasonality, holidays, and trends, and can incorporate additional regressors like promotions. This requires only SQL, matching the team's low-code requirement and the need to forecast per store.

Why this answer

BigQuery ML's ARIMA_PLUS is purpose-built for time series forecasting and can handle multiple related time series by specifying an identifier column. It automatically models seasonality, holidays, and trends, and supports additional regressors. This requires only SQL, making it ideal for teams with limited ML expertise who need per-store forecasts.

Exam trap

The trap here is assuming that any regression model can handle time series forecasting, but linear regression fails to capture temporal patterns like seasonality and trends.

27
MCQeasy

A company wants to classify customer support emails into categories like 'billing', 'technical', or 'account'. They have labeled email text data. Which AutoML solution should they use?

A.AutoML Tables
B.AutoML Natural Language
C.AutoML Video
D.AutoML Vision
AnswerB

AutoML Natural Language performs text classification on labelled data, training a model that assigns support emails to categories such as billing, technical or account, matching the labelled email text and multi-class requirement in the stem.

Why this answer

AutoML Natural Language is designed for text classification tasks, including sentiment analysis, entity extraction, and content categorization. Since the company has labeled email text data and wants to classify emails into categories like 'billing', 'technical', or 'account', AutoML Natural Language is the correct choice. It handles text data natively and provides a simple interface to train custom models without requiring deep ML expertise.

Exam trap

PMLE often tests the distinction between AutoML services based on data modality; candidates may confuse AutoML Natural Language with AutoML Tables when the data is text but stored in a tabular format.

How to eliminate wrong answers

Option A is wrong because AutoML Tables is for structured tabular data, not unstructured text. Option C is wrong because AutoML Video is for video classification, object tracking, and action recognition, not text. Option D is wrong because AutoML Vision is for image classification, object detection, and image segmentation, not text.

28
MCQmedium

A hospital's radiology department wants to build a model that flags possible pneumonia on chest X-rays. They have 8,000 labeled DICOM studies in a Cloud Storage bucket and no in-house data science staff. They need a managed service that can ingest the images, train a classifier, and provide an endpoint for their viewing software. What should they do?

A.Use Vertex AI Vision to build a pipeline that counts objects in the X-ray streams
B.Use the Cloud Vision API label detection feature on each X-ray
C.Create a Vertex AI dataset from the images and train an AutoML image classification model
D.Deploy a pre-trained TensorFlow Hub model to Vertex AI Endpoints without retraining
AnswerC

AutoML image classification accepts labeled images referenced from Cloud Storage, trains a managed classifier, and deploys an endpoint callable by the viewing software. It requires no model code, fitting a radiology team without data scientists who need a fast, managed path to a pneumonia flagger.

Why this answer

AutoML image classification is the managed, low-code route for a labeled image set: it reads images from Cloud Storage, trains and tunes a classifier, and serves predictions through an endpoint the viewing software can call. The alternatives either use fixed-label APIs, skip the essential fine-tuning, or apply a streaming analytics product instead of a diagnostic classifier.

Exam trap

The trap here is reaching for a general Vision API because it 'sees images,' when it cannot be trained on the department's pneumonia labels.

29
MCQeasy

A marketing team wants to build a model that predicts customer lifetime value (CLV) using historical transaction data. They are comfortable with spreadsheets but have no coding experience. They need a low-code solution that automatically handles feature engineering and model selection. Which Google Cloud service should they use?

A.AI Platform Training
B.Vertex AI AutoML
C.BigQuery ML
D.Vertex AI Workbench
AnswerB

Vertex AI AutoML provides a fully managed, no-code environment where users can upload data and automatically train models with feature engineering and hyperparameter tuning handled by the service. It supports tabular data for regression tasks like predicting CLV. The marketing team can use the UI without writing code, making it the ideal low-code solution.

Why this answer

Vertex AI AutoML is designed for users with limited ML expertise to build models without writing code. It automates feature engineering, model selection, and hyperparameter tuning. For predicting CLV from historical data, the marketing team can simply upload a dataset and let AutoML train a regression model.

Other options require SQL or Python coding, which the team lacks.

Exam trap

The trap here is assuming that BigQuery ML is a no-code solution because it uses SQL, but it still requires query writing and manual model configuration.

30
MCQmedium

A data scientist needs to forecast daily sales for the next 30 days using historical sales data stored in BigQuery. They want to use BigQuery ML. Which model type should they choose?

A.LINEAR_REG
B.BOOSTED_TREE_REGRESSOR
C.K_MEANS
D.ARIMA_PLUS
AnswerD

ARIMA_PLUS handles time-series forecasting natively in BigQuery ML, modelling trend, seasonality and holidays directly from historical data. It satisfies the requirement to forecast the next 30 days of daily sales without exporting data, and supports the forecast horizon through its horizon parameter.

Why this answer

ARIMA_PLUS is the correct choice because it is specifically designed for time-series forecasting, such as predicting daily sales over a future horizon. BigQuery ML's ARIMA_PLUS model automatically handles seasonality, trend, and holiday effects, making it ideal for 30-day sales forecasts from historical data.

Exam trap

The trap here is that candidates often confuse regression models (like LINEAR_REG or BOOSTED_TREE_REGRESSOR) with time-series forecasting, not realizing that standard regression assumes independent observations and cannot inherently model temporal dependencies or extrapolate beyond the training period.

How to eliminate wrong answers

Option A is wrong because LINEAR_REG is a linear regression model for predicting a continuous target from input features, but it does not inherently model time-series dependencies like autocorrelation or seasonality, making it unsuitable for forecasting sequential daily sales. Option B is wrong because BOOSTED_TREE_REGRESSOR is an ensemble tree-based model for regression tasks, but it treats each row independently and cannot capture temporal patterns or extrapolate into the future without explicit feature engineering of time lags. Option C is wrong because K_MEANS is an unsupervised clustering algorithm used to partition data into groups, not for forecasting numerical values over time.

31
MCQmedium

A data engineer wants to use BigQuery ML to train a model for predicting customer churn (binary classification) using a large dataset. They want the model to be automatically tuned. Which model type should they choose?

A.LOGISTIC_REG
B.BOOSTED_TREE_CLASSIFIER
C.DNN_CLASSIFIER
D.AUTOML_CLASSIFIER
AnswerD

AUTOML_CLASSIFIER satisfies the automatic tuning constraint: BigQuery ML performs hyperparameter tuning and architecture search internally, so no manual configuration is needed. It handles binary classification directly, matching the churn prediction task, unlike boosted tree or logistic regression models, which require the engineer to specify tuning themselves.

Why this answer

(AUTOML_CLASSIFIER) is correct because it automatically performs architecture search and hyperparameter tuning to find the best model for binary classification tasks, such as customer churn prediction. This is ideal when the data engineer wants the model to be automatically tuned without manual intervention, as AutoML handles feature engineering, model selection, and tuning under the hood.

Exam trap

The trap here is that candidates often confuse 'automatically tuned' with models that have default hyperparameters (like LOGISTIC_REG or BOOSTED_TREE_CLASSIFIER), but only AUTOML_CLASSIFIER performs automated hyperparameter tuning and architecture search without requiring manual specification.

How to eliminate wrong answers

Option A (LOGISTIC_REG) is wrong because logistic regression does not support automatic tuning; it requires manual specification of hyperparameters like learning rate or regularization, and it is a simpler linear model that may not capture complex patterns in large datasets. Option B (BOOSTED_TREE_CLASSIFIER) is wrong because while it can be tuned, it does not offer fully automatic tuning; the user must manually set parameters such as tree depth, learning rate, and number of iterations. Option C (DNN_CLASSIFIER) is wrong because deep neural network classifiers require manual tuning of architecture (e.g., number of layers, neurons) and hyperparameters (e.g., learning rate, batch size), and they do not automatically search for the optimal configuration.

32
MCQeasy

A data analyst wants to train a binary classification model on a BigQuery table without moving data out of BigQuery. They have limited ML expertise. Which approach should they take?

A.Use BigQuery ML with CREATE MODEL and LOGISTIC_REG model type.
B.Use Cloud Datalab to train an XGBoost model on BigQuery data.
C.Train a model using Vertex AI Workbench with a custom container.
D.Export the data to Cloud Storage and use Vertex AI AutoML Tables.
AnswerA

BigQuery ML trains models inside BigQuery using SQL, so data never leaves the warehouse. LOGISTIC_REG builds binary classification directly on the table, and CREATE MODEL keeps the workflow within SQL, matching the limited ML expertise and no-data-movement constraints.

Why this answer

BigQuery ML allows users to create and train binary classification models directly on data in BigQuery using SQL, with no need to move data or have deep ML expertise. The LOGISTIC_REG model type implements logistic regression, a standard algorithm for binary classification, and the CREATE MODEL statement handles all the underlying training infrastructure, making it ideal for a data analyst with limited ML skills.

Exam trap

Google often tests the distinction between low-code/no-code solutions (like BigQuery ML) and more advanced, infrastructure-heavy approaches (like custom containers or AutoML with data export), expecting candidates to recognize that the simplest, most integrated option is correct when the user has limited ML expertise and wants to avoid data movement.

How to eliminate wrong answers

Option B is wrong because Cloud Datalab is a deprecated interactive notebook service that requires users to write custom code and manage infrastructure, which is not suitable for someone with limited ML expertise and does not leverage BigQuery's native ML capabilities. Option C is wrong because Vertex AI Workbench with a custom container demands advanced knowledge of containerization, model training pipelines, and infrastructure management, far beyond the scope of a low-code solution for a data analyst. Option D is wrong because exporting data to Cloud Storage and using Vertex AI AutoML Tables, while low-code, introduces unnecessary data movement and additional complexity compared to the simpler, fully integrated BigQuery ML approach that keeps data in place.

33
MCQhard

A financial services company uses Document AI to process loan applications. They want to ensure that any documents the model cannot process with high confidence are reviewed by a human before finalizing the decision. Which Document AI feature should they enable?

A.AutoML Tables model retraining
B.Cloud DLP for data inspection
C.Increase the number of processors
D.Human-in-the-Loop (HITL)
AnswerD

Human-in-the-Loop routes low-confidence extractions to human reviewers before decisions finalise, directly satisfying the requirement that uncertain documents receive manual review. Document AI assigns confidence scores per field, and HITL triggers review when those scores fall below your configured threshold, ensuring no unreliable output reaches the loan decision.

Why this answer

Human-in-the-Loop (HITL) in Document AI allows you to route documents that the model processes with low confidence to human reviewers for validation or correction before finalizing. This directly matches the requirement to have a human review any documents the model cannot process with high confidence.

Exam trap

PMLE often tests the difference between model improvement features (retraining, AutoML) and operational review features (HITL) — candidates may choose retraining when the requirement is for a human review step, not model accuracy improvement.

How to eliminate wrong answers

Option A is wrong because AutoML Tables is for tabular data, not document processing, and retraining does not provide a human review workflow. Option B is wrong because Cloud DLP is for data inspection and redaction (PII discovery), not for human review of document extraction confidence. Option C is wrong because increasing the number of processors scales throughput but does not add a human review step for low-confidence documents.

34
MCQmedium

A financial services firm wants to predict loan default risk using a dataset with 30,000 labeled examples and 25 numeric and categorical features. Their team includes SQL analysts but no Python developers, and they want to minimize operational overhead. They decide to use BigQuery ML. Which model type should they use to achieve the best predictive performance while keeping the solution low-code?

A.BOOSTED_TREE_CLASSIFIER
B.DNN_CLASSIFIER
C.KMEANS
D.LOGISTIC_REG
AnswerA

BOOSTED_TREE_CLASSIFIER is an ensemble model that often achieves higher accuracy on tabular data with non-linear relationships, such as loan default prediction. It is fully supported in BigQuery ML, requires only SQL to train, and automatically handles feature engineering like categorical encoding. With 30,000 examples, it has sufficient data to train effectively without overfitting, making it ideal for this low-code, high-performance scenario.

Why this answer

For a binary classification task with a moderate-sized tabular dataset and a low-code requirement, BigQuery ML's BOOSTED_TREE_CLASSIFIER offers an excellent balance of accuracy and ease of use. It automatically handles categorical features and non-linear relationships, often outperforming linear models. It requires only SQL to train and deploy, aligning with the team's skills and minimizing operational overhead.

Exam trap

The trap here is assuming that a simple linear model like logistic regression is sufficient for all classification tasks, when actually boosted trees often yield better performance on tabular data without added complexity.

35
MCQmedium

A company needs to forecast product demand for the next 12 months using historical sales data. They want to use BigQuery ML with minimal coding. Which model type is most suitable?

A.K_MEANS
B.MATRIX_FACTORIZATION
C.ARIMA_PLUS
D.LINEAR_REG
AnswerC

ARIMA_PLUS handles time-series forecasting natively in BigQuery ML, requiring only a single CREATE MODEL statement on the historical sales column. It automatically detects seasonality, trends and holidays, satisfying the 12-month horizon and minimal-coding constraint without exporting data or writing Python.

Why this answer

ARIMA_PLUS is BigQuery ML's purpose-built time-series forecasting model, designed for exactly this scenario: forecasting future values from historical time-ordered data with minimal SQL coding. It automatically handles seasonality, holidays, and trend decomposition, and supports features like forecasting multiple time series at once. LINEAR_REG could technically model time as a feature, but it cannot capture seasonality or autocorrelation, making it unsuitable for demand forecasting.

Exam trap

The trap here is that candidates see 'forecast' and reach for LINEAR_REG because it is a regression model, forgetting that time-series forecasting requires specialized models like ARIMA_PLUS that handle seasonality and autocorrelation.

How to eliminate wrong answers

Option A is wrong because K_MEANS is an unsupervised clustering algorithm used for segmentation, not for predicting future numeric values over time. Option B is wrong because MATRIX_FACTORIZATION is a collaborative-filtering technique for recommendation systems (e.g., user-item matrices), not time-series forecasting. Option D is wrong because LINEAR_REG assumes independent observations and cannot model seasonality, trends, or autocorrelated time-series structure, so it would produce poor demand forecasts.

36
MCQeasy

A developer wants to add text translation to a mobile app. They need to translate user-generated content into multiple languages, and latency is critical. Which pre-built API should they use?

A.Translation API
B.Vision API
C.Text-to-Speech API
D.Natural Language API
AnswerA

The Translation API is a pre-built, low-latency service supporting many language pairs, so it meets the requirement to translate user-generated content into multiple languages without training custom models. Custom translation would add latency and development overhead.

Why this answer

Translation API provides fast, real-time translation for text. Natural Language API is for analysis. Text-to-Speech is for audio.

Vision API is for images.

37
Multi-Selectmedium

A company wants to build a model to predict housing prices using BigQuery ML. They have a dataset with features like area, number of bedrooms, and location. Which TWO model types are appropriate for this regression task?

Select 2 answers
A.LOGISTIC_REG
B.K_MEANS
C.MATRIX_FACTORIZATION
D.BOOSTED_TREE_REGRESSOR
E.LINEAR_REG
AnswersD, E

Why this answer

BOOSTED_TREE_REGRESSOR (D) is appropriate because it is a tree-based ensemble method specifically designed for regression tasks, and BigQuery ML supports it via the `CREATE MODEL` statement with `model_type='BOOSTED_TREE_REGRESSOR'`. It handles non-linear relationships and interactions between features like area, bedrooms, and location, making it suitable for predicting continuous housing prices.

Exam trap

The PMLE exam often tests the distinction between regression and classification models, leading candidates to mistakenly choose LOGISTIC_REG for regression tasks because of the word 'regression' in its name, but it is actually a classification algorithm.

38
MCQeasy

A marketing team wants to build a model that predicts whether a customer will click on an ad, using a dataset in BigQuery. They have limited ML expertise and want to avoid writing complex code. They decide to use BigQuery ML with a logistic regression model. Which SQL statement should they use to create the model?

A.CREATE MODEL `project.dataset.model` OPTIONS(model_type='logistic_reg', input_label_cols=['clicked']) AS SELECT * FROM `project.dataset.training_data`
B.CREATE MODEL `project.dataset.model` OPTIONS(model_type='linear_reg', input_label_cols=['clicked']) AS SELECT * FROM `project.dataset.training_data`
C.CREATE MODEL `project.dataset.model` OPTIONS(model_type='logistic_reg') AS SELECT * FROM `project.dataset.training_data`
D.CREATE MODEL `project.dataset.model` OPTIONS(model_type='kmeans', input_label_cols=['clicked']) AS SELECT * FROM `project.dataset.training_data`
AnswerA

This statement correctly creates a logistic regression model in BigQuery ML. It specifies the model type as logistic_reg and identifies the label column as 'clicked' using the input_label_cols option. The SELECT statement provides the training data, which includes both features and the label. BigQuery ML will automatically use all other columns as features, making it suitable for users with limited ML expertise.

Why this answer

The correct SQL statement creates a logistic regression model in BigQuery ML with the proper model type and label column specification. Logistic regression is ideal for binary classification tasks such as predicting ad clicks. The input_label_cols option identifies the target variable, and the SELECT statement supplies the training data.

This approach allows users with limited ML expertise to build a model using only SQL.

Exam trap

The trap here is confusing linear regression with logistic regression for binary classification tasks.

39
MCQhard

A financial institution wants to detect fraudulent transactions in real-time. They have a labeled dataset of historical transactions and want to use a low-code solution that can automatically handle feature engineering and model selection. They also need to deploy the model for online predictions with low latency. Which Google Cloud service should they use?

A.Vertex AI AutoML Tables
B.Vertex AI Pipelines with Kubeflow
C.Cloud Functions with a custom Python script
D.BigQuery ML with a logistic regression model
AnswerA

Vertex AI AutoML Tables automatically performs feature engineering, including one-hot encoding, normalization, and feature selection, and selects the best model architecture. It supports online prediction with low latency when deployed to an endpoint. This aligns perfectly with the need for a low-code solution that handles feature engineering and provides real-time fraud detection. It is the most suitable choice for this scenario.

Why this answer

Vertex AI AutoML Tables is the correct choice because it provides automatic feature engineering and model selection, and supports low-latency online predictions when deployed to an endpoint. It is designed for tabular data like financial transactions and requires minimal coding. Other options either lack automatic feature engineering, require significant coding, or are not intended for building predictive models, making AutoML Tables the best fit for real-time fraud detection.

Exam trap

The trap here is assuming that BigQuery ML models can be directly used for low-latency online predictions without additional deployment steps, when in fact Vertex AI AutoML Tables offers a more streamlined path.

40
MCQeasy

A financial services company wants to extract text and structured data from scanned loan application forms. They need a fully managed, low-code solution that can handle various form layouts and requires minimal machine learning expertise. Which Google Cloud service should they use?

A.Document AI
B.Vision API
C.BigQuery ML
D.AutoML Natural Language
AnswerA

Document AI is a fully managed service that uses pre-trained and customizable models to parse documents, extract text, and identify structured fields. It supports form parsing and can be used with minimal ML expertise via the console or API. It is ideal for processing loan applications with varying layouts.

Why this answer

Document AI provides specialized document parsing capabilities, including form understanding and entity extraction, with pre-built models and the ability to train custom extractors with minimal effort. It is fully managed and designed for low-code document processing, making it the best fit for loan application forms.

Exam trap

The trap here is assuming that Vision API's OCR is sufficient, but it lacks structured extraction and form parsing capabilities.

41
MCQmedium

A company has a large dataset of labeled images (e.g., different species of plants). They want to train a custom image classification model with minimal effort and no prior ML experience. Which Google Cloud service should they use?

A.Cloud TPU
B.AutoML Vision
C.Vertex AI Workbench with a custom TensorFlow model
D.Vision API
AnswerB

AutoML Vision trains custom image classifiers through a point-and-click interface, using transfer learning on Google's pretrained models. This satisfies the stem's constraints: labelled plant images, minimal effort, and no ML expertise. Vertex AI Vision now supersedes it, but AutoML Vision remains the purpose-built answer for codeless custom classification.

Why this answer

AutoML Vision is the correct choice because it allows users with no prior ML experience to train a custom image classification model using a simple graphical interface, requiring only labeled images as input. It automates model architecture search, hyperparameter tuning, and deployment, minimizing manual effort while delivering a production-ready model.

Exam trap

The trap here is that candidates confuse AutoML Vision (custom model training with minimal effort) with Vision API (pre-trained, no custom training), often picking D because both involve 'Vision' and seem low-code, but Vision API cannot be retrained on custom data.

How to eliminate wrong answers

Option A is wrong because Cloud TPU is a hardware accelerator for training custom models, requiring users to write and manage their own ML code (e.g., TensorFlow/PyTorch), which demands significant ML expertise and effort. Option C is wrong because Vertex AI Workbench with a custom TensorFlow model requires users to write, debug, and train a model from scratch using notebooks, which is not minimal effort and assumes ML experience. Option D is wrong because Vision API is a pre-trained API for general image recognition tasks (e.g., label detection, OCR) and cannot be trained on custom labeled datasets like plant species; it offers no customization for specific classification needs.

42
MCQeasy

A company wants to transcribe customer service calls in real-time to detect sentiment and identify urgent issues. They need a solution with low latency. Which combination of pre-built APIs should they use?

A.Text-to-Speech and Natural Language API
B.Speech-to-Text and Translation API
C.Video Intelligence API
D.Speech-to-Text and Natural Language API
AnswerD

Speech-to-Text streams audio into text with low latency, and the Natural Language API then performs sentiment analysis and entity detection on that text. Together they deliver real-time transcription plus sentiment and urgency identification without training custom models.

Why this answer

Real-time call transcription requires Speech-to-Text to convert audio to text with low latency, and Natural Language API to analyze that text for sentiment and entity/urgency detection. Together they form the standard streaming audio-analysis pipeline on Google Cloud. Speech-to-Text supports streaming recognition, and Natural Language provides sentiment and entity analysis on the resulting transcript.

Exam trap

PMLE often tests whether candidates confuse the direction of the audio APIs, so picking Text-to-Speech (which synthesizes speech) instead of Speech-to-Text (which transcribes) is the classic wrong-answer trap for transcription scenarios.

How to eliminate wrong answers

Option A is wrong because Text-to-Speech converts text to audio — the opposite direction — and cannot transcribe incoming calls. Option B is wrong because Translation API translates between languages; it does not perform sentiment analysis or urgency detection, so it fails the stated requirement. Option C is wrong because Video Intelligence API analyzes video content (labels, shots, explicit content) and is not designed for real-time audio transcription or sentiment analysis.

43
MCQhard

A financial services firm wants relationship managers to query a natural-language interface such as 'show me customers likely to churn next quarter' and receive results from their BigQuery data warehouse. They want to minimize the code they maintain and keep data in BigQuery. Which approach best fits?

A.Enable Gemini in BigQuery to generate SQL from natural-language questions over the warehouse
B.Export the warehouse to Cloud Storage and train a Vertex AI AutoML Tables model nightly
C.Build a custom chatbot with Dialogflow CX and hand-written SQL for each intent
D.Use BigQuery ML to train an ARIMA_PLUS model on customer activity to predict churn
AnswerA

Gemini in BigQuery translates natural-language prompts into SQL against the firm's tables and schemas, letting managers ask churn-style questions without writing queries. It keeps data in BigQuery and requires no bespoke codebase, matching the low-maintenance, in-warehouse requirement.

Why this answer

Gemini in BigQuery converts natural-language prompts into SQL executed against the firm's existing tables, so relationship managers get ad hoc answers without maintaining query code and data never leaves BigQuery. The alternatives either require hand-written SQL per intent, apply an unsuitable forecasting model, or move data out of the warehouse into a batch training pipeline.

Exam trap

The trap here is equating a conversational agent like Dialogflow CX with natural-language-to-SQL, when the former needs predefined intents and queries.

44
MCQmedium

A company wants to build a product recommendation engine for their e-commerce website. They have historical purchase data and user interaction logs. They want a managed service that can quickly generate personalized recommendations without building custom models. Which service should they use?

A.Dataflow with TensorFlow
B.BigQuery ML with MATRIX_FACTORIZATION
C.AutoML Tables
D.Recommendations AI
AnswerD

Recommendations AI is a managed Google Cloud service that trains on historical purchase and interaction data to serve personalised product recommendations, satisfying the requirement to avoid building custom models. It handles model training, tuning and serving, so the team gains personalisation without ML engineering effort.

Why this answer

Recommendations AI is a fully managed Google Cloud service purpose-built for personalized product recommendations, ingesting purchase and interaction data and serving recommendations via API without custom model development. It handles training, tuning, and serving, which matches the requirement for a managed service that quickly generates personalized recommendations.

Exam trap

PMLE often tests the confusion between general ML services (AutoML, BigQuery ML) and purpose-built recommendation services, where only the latter meets a 'no custom models' requirement.

How to eliminate wrong answers

Option A is wrong because Dataflow with TensorFlow requires building and operating a custom recommendation model, which contradicts the 'without building custom models' requirement. Option B is wrong because BigQuery ML with MATRIX_FACTORIZATION requires writing SQL to train and tune a model and does not provide a turnkey recommendation serving API. Option C is wrong because AutoML Tables is a general tabular ML service, not a recommendation-specific product, and would still require significant feature engineering and serving infrastructure.

45
MCQhard

A company uses BigQuery ML to train a boosted tree classifier on a large dataset. After training, they want to understand which features most influence predictions. Which BigQuery ML function should they use?

A.ML.EXPLAIN_PREDICT
B.ML.FEATURE_IMPORTANCE
C.ML.EVALUATE
D.ML.PREDICT
AnswerB

ML.FEATURE_IMPORTANCE returns a per-feature score showing how strongly each input influenced the trained boosted tree model's predictions. It is the BigQuery ML function designed for interpreting model behaviour, satisfying the requirement to identify the most influential features.

Why this answer

ML.FEATURE_IMPORTANCE is the BigQuery ML function specifically designed to return the relative importance of each input feature for a trained model, including boosted tree classifiers (Boosted Tree, XGBoost, Random Forest). It computes importance scores using the model's internal split-gain or weight-based metrics, giving a ranked list of features that most influence predictions. This is the direct, purpose-built answer for global feature attribution on a trained model.

Exam trap

PMLE often tests the confusion between global feature importance (ML.FEATURE_IMPORTANCE) and local per-prediction explanations (ML.EXPLAIN_PREDICT) — candidates pick EXPLAIN_PREDICT because it sounds more 'explanatory,' but it returns Shapley values per row, not a global ranking.

How to eliminate wrong answers

Option A is wrong because ML.EXPLAIN_PREDICT returns per-row local explanations (Shapley values) for individual predictions, not a global ranked feature-importance list — it answers 'why this prediction' rather than 'which features matter overall.' Option C is wrong because ML.EVALUATE returns model quality metrics (accuracy, precision, recall, AUC, log loss) on a dataset, not feature attribution. Option D is wrong because ML.PREDICT generates predictions on new data; it provides no insight into feature influence.

46
MCQeasy

A retail company wants to build a recommendation system to show 'frequently bought together' items. Which Recommendations AI model type should they use?

A.recently-viewed
B.frequently-bought-together
C.recommended-for-you
D.others-you-may-like
AnswerB

The frequently-bought-together model type is purpose-built to surface item pairs or sets commonly purchased in the same transaction, matching the retail 'frequently bought together' requirement. Other model types target different goals such as personalised ranking or similar-item recommendations.

Why this answer

The 'frequently-bought-together' model type in Google Cloud Recommendations AI is specifically designed to identify items that are commonly purchased in the same transaction, using co-occurrence analysis of historical purchase data. This directly matches the requirement to show items that are frequently bought together, leveraging association rule mining (e.g., Apriori algorithm) to generate recommendations.

Exam trap

In Google Cloud Recommendations AI, the key distinction is between transaction-based co-purchase models ('frequently-bought-together') and personalized recommendation models ('recommended-for-you'). Candidates may confuse 'frequently-bought-together' with 'others-you-may-like' because both relate to item similarity, but only 'frequently-bought-together' uses co-occurrence analysis of actual transactions.

How to eliminate wrong answers

Option A is wrong because 'recently-viewed' is a model type that surfaces items a user has recently browsed, not items that are frequently purchased together, and it relies on session-based user behavior rather than transaction co-occurrence. Option C is wrong because 'recommended-for-you' is a personalized model that uses user-item interaction history (e.g., collaborative filtering) to suggest items tailored to an individual, not cross-item purchase patterns. Option D is wrong because 'others-you-may-like' is a similarity-based model that recommends items similar to a given product based on content or metadata, not on transactional co-purchase frequency.

47
MCQhard

An engineer wants to use BigQuery ML to explain predictions from a trained boosted tree classifier for a specific set of input rows. Which function should they use?

A.ML.EVALUATE
B.ML.FEATURE_IMPORTANCE
C.ML.PREDICT
D.ML.EXPLAIN_PREDICT
AnswerD

ML.EXPLAIN_PREDICT returns predictions alongside feature attributions for the supplied rows, using explainable AI methods such as Shapley values or integrated gradients, which is exactly the per-row explanation the engineer needs from the boosted tree classifier.

Why this answer

ML.EXPLAIN_PREDICT is the BigQuery ML function specifically designed to explain predictions from a trained model for a given set of input rows. It returns the prediction along with feature attributions (e.g., Shapley values) that show how each feature contributed to the prediction. This is the correct function for interpreting individual predictions from a boosted tree classifier.

Exam trap

PMLE often tests the confusion between global explainability (ML.FEATURE_IMPORTANCE, ML.GLOBAL_EXPLAIN) and local explainability (ML.EXPLAIN_PREDICT) — candidates must know which function provides per-row feature attributions.

How to eliminate wrong answers

Option A is wrong because ML.EVALUATE computes overall model metrics (accuracy, precision, recall) on a dataset, not explanations for individual predictions. Option B is wrong because ML.FEATURE_IMPORTANCE returns global feature importance for the model, not per-row explanations. Option C is wrong because ML.PREDICT returns predictions without explanations; it does not provide feature attributions.

48
Multi-Selecthard

A company is building a document processing pipeline for invoices. They need to extract key fields (invoice number, date, total amount) and allow human review for invoices over $10,000. Which TWO Google Cloud services/features should they combine?

Select 2 answers
A.Cloud Vision API for OCR
B.Human-in-the-Loop (HITL) on Document AI
C.AutoML Tables to predict missing fields
D.Document AI with invoice parser processor
E.Cloud Translation API to translate invoices
AnswersB, D

Why this answer

Option D is correct because Document AI's specialized Invoice parser processor is purpose-built to extract structured fields such as invoice number, invoice date, and total amount from invoice documents, which is exactly the extraction requirement in this pipeline. Option B is correct because Document AI's Human-in-the-Loop (HITL) feature lets you define confidence thresholds and route documents for manual review, which directly satisfies the requirement to have humans review invoices over $10,000 before downstream processing. Together, the Invoice parser handles automated field extraction while HITL provides the human verification step for high-value invoices.

Option A is not the best fit because Cloud Vision API provides generic OCR without invoice-specific field extraction or built-in human review workflows. Option C is incorrect because AutoML Tables is a tabular ML service for structured data prediction, not for predicting missing fields in parsed documents. Option E is incorrect because Cloud Translation API only translates text and does not extract invoice fields or support human review.

Exam trap

PMLE often tests whether candidates pick generic OCR (Cloud Vision) instead of the domain-specific Document AI parser, and whether they recognize HITL as the mechanism for human review rather than building custom review logic.

49
MCQeasy

A data analyst wants to train a linear regression model to predict house prices using only SQL queries on BigQuery. Which BigQuery ML model type should they use?

A.BOOSTED_TREE_REGRESSOR
B.LOGISTIC_REG
C.LINEAR_REG
D.DNN_REGRESSOR
AnswerC

LINEAR_REG is BigQuery ML's built-in linear regression model type, trained directly with CREATE MODEL on a numeric label such as house price. It satisfies the constraint of using only SQL, requiring no exported data or external framework.

Why this answer

The question specifies a linear regression model for predicting house prices, which is a regression task with a continuous target variable. BigQuery ML's LINEAR_REG model type is explicitly designed for linear regression, making it the correct choice for this use case.

Exam trap

Google often tests the distinction between regression and classification model types, and the trap here is that candidates might confuse LOGISTIC_REG (classification) with linear regression due to the word 'logistic' sounding similar to 'linear', or they might overcomplicate the solution by choosing a tree or neural network model when a simple linear model suffices.

How to eliminate wrong answers

Option A is wrong because BOOSTED_TREE_REGRESSOR is a tree-based ensemble method, not a linear model, and is overkill for a simple linear regression task. Option B is wrong because LOGISTIC_REG is used for binary classification, not regression (predicting continuous values like house prices). Option D is wrong because DNN_REGRESSOR is a deep neural network regressor, which is unnecessarily complex and not a linear model.

50
MCQmedium

A developer needs to transcribe phone calls with high accuracy for a call center analytics application. The audio is in English and has background noise. Which Speech-to-Text model should they choose?

A.telephony
B.latest_short
C.latest_long
D.Any model; they are all equivalent
AnswerA

The telephony model is trained on 8 kHz narrowband audio, matching the bandwidth of phone calls, and is optimised to suppress background noise. This directly satisfies the stem's call centre constraint: English telephone speech with noise, where the default model's wideband training would degrade accuracy.

Why this answer

The 'telephony' model is specifically optimized for transcribing audio from phone calls, which often includes background noise, low fidelity, and narrowband audio. It is designed to handle the unique characteristics of telephony audio, such as 8kHz sampling rate and compression artifacts, providing higher accuracy for call center analytics. Other models like 'latest_short' and 'latest_long' are general-purpose and may not perform as well on noisy phone calls.

Exam trap

PMLE often tests the misconception that any Speech-to-Text model can be used interchangeably, but the trap is failing to recognize that telephony audio requires a specialized model due to its unique acoustic properties.

How to eliminate wrong answers

Option B is wrong because 'latest_short' is optimized for short utterances (e.g., voice commands) and does not handle the background noise and telephony-specific audio characteristics as effectively. Option C is wrong because 'latest_long' is designed for long-form audio like interviews or lectures, but it is not specifically tuned for telephony noise and compression. Option D is wrong because models are not equivalent; each is optimized for different audio types, and using the wrong model can significantly reduce transcription accuracy.

51
MCQmedium

A data scientist wants to use AutoML to classify images of retail products into categories. There are 50 categories and the dataset has 100,000 labelled images. Which Vertex AI AutoML service is most appropriate?

A.AutoML Tables
B.AutoML Video
C.AutoML Vision
D.AutoML NLP
AnswerC

AutoML Vision handles multi-class image classification directly, supporting up to 50 categories without custom architecture work. It ingests the 100,000 labelled images and trains a managed model, satisfying the stem's classification requirement. Vertex AI's image data type is the appropriate task-specific service here, unlike tabular or video alternatives.

Why this answer

AutoML Vision is Vertex AI's service for training custom image classification and object detection models. With 100,000 labeled images across 50 categories, AutoML Vision can train a multi-class image classifier that learns visual features for each product category. It is the appropriate choice when the input data is images and the task is classification.

Exam trap

PMLE often tests whether candidates match the data modality (image vs. tabular vs. text vs. video) to the correct AutoML service, so choosing AutoML Tables for image data is the common error.

How to eliminate wrong answers

Option A is wrong because AutoML Tables is for structured/tabular data (CSV, BigQuery), not images. Option B is wrong because AutoML Video is for video classification, object tracking, and action recognition, not still-image product classification. Option D is wrong because AutoML NLP handles text classification, entity extraction, and sentiment analysis, not image data.

52
MCQeasy

A data scientist needs to train a time-series forecasting model on historical sales data stored in BigQuery to predict future demand. The data has strong seasonal patterns. Which BigQuery ML model type should they use?

A.MATRIX_FACTORIZATION
B.BOOSTED_TREE_REGRESSOR
C.ARIMA_PLUS
D.K_MEANS
AnswerC

ARIMA_PLUS handles seasonality natively through automatic seasonal decomposition and multiple seasonal period detection, satisfying the strong seasonal patterns constraint in the historical sales data. It also supports forecasting horizons directly in BigQuery ML, letting the data scientist train and predict demand without exporting data.

Why this answer

ARIMA_PLUS is the correct choice because it is specifically designed for time-series forecasting in BigQuery ML, handling seasonal patterns, trend decomposition, and automatic hyperparameter tuning. It models autoregressive (AR) and moving average (MA) components with seasonal differencing, making it ideal for historical sales data with strong seasonal cycles.

Exam trap

Google often tests the misconception that any regression model (like BOOSTED_TREE_REGRESSOR) can be naively applied to time-series data, ignoring the need for specialized models that handle temporal dependencies and seasonality natively.

How to eliminate wrong answers

Option A is wrong because MATRIX_FACTORIZATION is used for recommendation systems (e.g., collaborative filtering) and cannot model temporal dependencies or seasonality in time-series data. Option B is wrong because BOOSTED_TREE_REGRESSOR is a tree-based ensemble method for regression tasks but does not inherently capture time-series structures like seasonality, trend, or autocorrelation without extensive feature engineering. Option D is wrong because K_MEANS is an unsupervised clustering algorithm that groups data points by similarity and has no mechanism for forecasting future values or modeling sequential patterns.

53
MCQhard

A financial institution wants to detect fraudulent transactions in real-time. They have a labeled dataset of historical transactions and want to build a custom model with minimal coding. They also need to integrate the model into an existing application that expects a REST API. Which Google Cloud service should they use to train and deploy the model with the least effort?

A.Vertex AI AutoML Tabular to train a model, then deploy it to a Vertex AI endpoint.
B.AI Platform Training with a custom TensorFlow model, then deploy to AI Platform Prediction.
C.Cloud Functions to call the BigQuery ML model directly via a REST API.
D.BigQuery ML to train a logistic regression model, then export the model to Vertex AI for deployment.
AnswerA

Vertex AI AutoML Tabular automates model training and tuning for tabular data, minimizing coding. Once trained, the model can be deployed to a Vertex AI endpoint, which provides a REST API for real-time predictions. This end-to-end integration requires minimal effort and meets the requirement for a custom model with low-code and REST API access.

Why this answer

Vertex AI AutoML Tabular provides a low-code solution for training custom models on tabular data. It automatically handles feature engineering and model selection, and the trained model can be deployed to a Vertex AI endpoint, which offers a REST API for real-time predictions. This integration minimizes effort and meets the need for a custom model with REST API access.

Other options either require more coding, add unnecessary complexity, or do not support real-time serving natively.

Exam trap

The trap here is assuming that BigQuery ML can directly serve real-time REST API predictions, but it is primarily for batch or SQL-based predictions, not low-latency online serving.

54
MCQhard

A company wants to build a recommendation system that suggests products to users based on their past interactions. They have user-item interaction data in BigQuery and want a low-code solution that can generate recommendations for all users. Which approach should they use?

A.Use BigQuery ML to train a k-means clustering model and assign users to clusters for recommendations.
B.Use Dataflow to preprocess data and train a custom recommendation model on Vertex AI.
C.Use Vertex AI AutoML Tables to train a model that predicts a rating for each user-item pair.
D.Use BigQuery ML to train a matrix factorization model and use ML.RECOMMEND to generate recommendations.
AnswerD

BigQuery ML's matrix factorization model is specifically designed for recommendation tasks. It can be trained directly on user-item interaction data using SQL, which is low-code. The ML.RECOMMEND function then generates top-N recommendations for all users efficiently. This approach meets the requirement for a low-code solution that scales to all users, making it the ideal choice for this scenario.

Why this answer

BigQuery ML's matrix factorization model is purpose-built for recommendations and can be trained with SQL on user-item interaction data. The ML.RECOMMEND function generates recommendations for all users in a scalable way. This low-code approach avoids custom coding and is optimized for the task, making it the best fit for the company's needs.

Exam trap

The trap here is assuming that any BigQuery ML model can generate recommendations, when only matrix factorization and a few other specialized types support ML.RECOMMEND.

55
MCQmedium

A media company wants to automatically transcribe and analyze customer support calls to identify common issues. They need a low-code solution that provides both transcription and sentiment analysis. Which Google Cloud service should they use?

A.Speech-to-Text API
B.Contact Center AI Insights
C.Dialogflow CX
D.Natural Language API
AnswerB

Contact Center AI Insights is designed to analyze customer interactions, providing transcription, sentiment analysis, and topic detection out of the box. It is a low-code solution that integrates with various contact center platforms. It can automatically process calls and surface insights without custom model development. This directly meets the requirement for both transcription and sentiment analysis in a low-code manner.

Why this answer

Contact Center AI Insights is a purpose-built, low-code solution for analyzing customer calls. It automatically transcribes audio, performs sentiment analysis, and extracts topics, providing actionable insights without custom development. Other options either lack sentiment analysis or require integration with multiple services, increasing complexity and coding effort.

Thus, Contact Center AI Insights is the correct choice.

Exam trap

The trap here is assuming that combining Speech-to-Text and Natural Language API is low-code; it actually requires integration work, whereas Contact Center AI Insights provides an all-in-one managed solution.

56
MCQmedium

A retail company wants to predict customer churn using historical purchase data stored in BigQuery. The data includes customer demographics, transaction history, and support interactions. The team is comfortable writing SQL and wants to avoid moving data to a separate environment. Which approach should they take?

A.Use the Cloud Natural Language API to analyze customer support interactions and combine results with purchase data in BigQuery.
B.Export the data to a CSV file and use Vertex AI AutoML Tables to train a classification model.
C.Use BigQuery ML to create a logistic regression model (LOGISTIC_REG) on the data directly in BigQuery.
D.Create a Dataflow pipeline to stream data to Cloud SQL and use Cloud SQL's built-in ML functions.
AnswerC

BigQuery ML trains LOGISTIC_REG models using SQL directly against BigQuery-resident data, so no extraction or separate environment is needed. This satisfies the team's SQL comfort and the constraint of avoiding data movement for churn prediction.

Why this answer

BigQuery ML allows the team to build and train a logistic regression model directly on data stored in BigQuery using SQL syntax, without moving data to a separate environment. The LOGISTIC_REG model type is specifically designed for binary classification tasks like churn prediction, and it runs entirely within BigQuery's serverless infrastructure, satisfying the team's requirement to avoid data movement.

Exam trap

This question tests the misconception that ML requires moving data to a separate platform (like Vertex AI or Cloud SQL), when in fact BigQuery ML provides a low-code, SQL-based solution that keeps data in place and meets the stated constraints.

How to eliminate wrong answers

Option A is wrong because the Cloud Natural Language API is used for text analysis (e.g., sentiment extraction), not for training a predictive churn model; it would require additional steps to combine results and does not provide a built-in classification model. Option B is wrong because exporting data to a CSV file and using Vertex AI AutoML Tables violates the requirement to avoid moving data to a separate environment, and it introduces unnecessary data egress and manual steps. Option D is wrong because Cloud SQL does not have built-in ML functions for training classification models; it is a relational database service, and streaming data through Dataflow to Cloud SQL adds complexity and does not leverage BigQuery's native ML capabilities.

57
MCQeasy

A retail company wants to forecast daily sales for the next 30 days based on historical sales data. They have two years of daily sales records with no missing values. They want to use a low-code solution on Google Cloud that automatically handles seasonality and trends. Which service should they use?

A.Cloud Dataflow with a sliding window aggregation
B.Vertex AI Forecasting with a custom training container
C.BigQuery ML with the ARIMA_PLUS model
D.AutoML Tables with a regression target
AnswerC

BigQuery ML's ARIMA_PLUS model is specifically designed for time series forecasting. It automatically handles seasonality, trends, and holiday effects, and can generate forecasts for multiple time steps. It requires only SQL to train and predict, making it a low-code solution. Given the clean daily sales data, ARIMA_PLUS is well-suited to produce accurate 30-day forecasts without manual feature engineering.

Why this answer

BigQuery ML's ARIMA_PLUS is the correct choice because it is a low-code, SQL-based time series forecasting model that automatically handles seasonality, trends, and holidays. It is designed for exactly this type of scenario: forecasting future values from historical time series data. Other options either require manual feature engineering, are not forecasting-specific, or involve significant coding, making ARIMA_PLUS the most efficient and appropriate solution.

Exam trap

The trap here is confusing general-purpose regression models with specialized time series models, overlooking that only ARIMA_PLUS automatically handles temporal patterns like seasonality.

58
MCQmedium

A healthcare provider needs to extract structured information from incoming PDF forms (e.g., patient intake forms). They want to automate data extraction without writing custom models. Which Google Cloud service should they use?

A.Document AI with a form parser processor
B.Natural Language API for entity extraction
C.Vision API
D.AutoML Vision for object detection
AnswerA

Document AI's form parser processor is a pretrained model that extracts key-value pairs and tables from PDFs, requiring no custom model training. This directly satisfies the constraint of automating structured extraction from intake forms without writing custom models.

Why this answer

Document AI with a form parser processor is the correct choice because it is purpose-built for extracting structured data from PDF forms, including key-value pairs and tables, without requiring custom model development. It uses pre-trained models specifically for form understanding, making it ideal for automating intake form processing.

Exam trap

A common pitfall is choosing the Vision API for form parsing because it performs OCR, but it lacks the specialized form-field extraction capabilities of Document AI, which is designed specifically for structured data extraction from forms.

How to eliminate wrong answers

Option B is wrong because Natural Language API is designed for extracting entities and sentiment from unstructured text, not for parsing structured form fields from PDF documents. Option C is wrong because Vision API provides optical character recognition (OCR) to extract raw text from images but lacks the form-specific parsing logic to identify key-value pairs and table structures. Option D is wrong because AutoML Vision for object detection is used for identifying objects within images, not for extracting structured data from forms.

59
MCQmedium

A retail company wants to forecast monthly sales for each of its 500 stores using historical sales data. They have two years of daily sales data per store and want to use BigQuery ML to build a forecasting model. They need to account for seasonality and trends. Which BigQuery ML model type should they use?

A.Linear regression
B.Logistic regression
C.K-means clustering
D.ARIMA_PLUS
AnswerD

ARIMA_PLUS is a built-in BigQuery ML model for time series forecasting that automatically handles seasonality, trends, and holidays. It is designed for forecasting multiple time series, such as sales per store, and can process large datasets. It provides explainable forecasts and requires only the time series identifier, timestamp, and value columns. This makes it ideal for the retail company's monthly sales forecasting across many stores.

Why this answer

ARIMA_PLUS is specifically designed for time series forecasting in BigQuery ML. It automatically detects and models seasonality, trends, and holidays, which are critical for retail sales data. It supports multiple time series, allowing per-store forecasts.

This low-code solution requires minimal feature engineering and provides accurate, explainable results, making it the right choice for the company's needs.

Exam trap

The trap here is using a general regression model like linear regression for time series data without accounting for seasonality.

60
MCQmedium

A retail company wants to generate product recommendations on their website using Google Cloud. They have historical transaction data and need a managed service that provides personalized recommendations like 'frequently bought together'. Which service should they use?

A.Recommendations AI
B.BigQuery ML
C.Vertex AI Prediction
D.AutoML Tables
AnswerA

Recommendations AI is a managed Google Cloud service that ingests historical transaction data and produces personalised product recommendations such as 'frequently bought together', satisfying the managed-service and recommendation-quality constraints. It removes the need to build and train custom recommendation models.

Why this answer

Recommendations AI is a purpose-built managed service on Google Cloud designed for retail recommendation use cases, including 'frequently bought together', 'recommended for you', and 'others you may like'. It ingests historical transaction and catalog data, trains retail-specific models automatically, and serves personalized recommendations via API without requiring ML expertise. This directly matches the requirement for a managed service providing personalized product recommendations.

Exam trap

The trap here is confusing general-purpose ML services (BigQuery ML, Vertex AI, AutoML) with a purpose-built managed recommendation service — candidates who default to 'Vertex AI for everything ML' miss that Recommendations AI is the retail-specific managed offering.

How to eliminate wrong answers

Option B is wrong because BigQuery ML lets you build and run ML models using SQL inside BigQuery, but it is not a managed recommendation service — you would have to design, train, and serve the recommender yourself. Option C is wrong because Vertex AI Prediction is only a model-serving endpoint; it hosts models but does not provide retail recommendation logic, training pipelines, or pre-built recommendation types. Option D is wrong because AutoML Tables is a general tabular AutoML trainer, not a recommendation-specific managed service, and it lacks the retail recommendation objectives and serving API that Recommendations AI provides.

61
MCQmedium

A marketing team needs to build a model that predicts whether a customer will respond to a promotional email. They have a BigQuery table with 2 million rows and 30 features, and they want to avoid writing any Python code. They require an explainable model and the ability to generate predictions directly in SQL. Which approach should they use?

A.Use BigQuery ML to create a k-means clustering model to segment customers, then manually assign response probabilities based on cluster averages.
B.Create a BigQuery ML logistic regression model using CREATE MODEL with model_type='LOGISTIC_REG', then use ML.PREDICT for scoring.
C.Export the data to Cloud Storage and train a custom TensorFlow model using Vertex AI Training with a pre-built container.
D.Use AutoML Tables (now Vertex AI Tabular) to train a model, then deploy it to an endpoint and call it from BigQuery using a remote function.
AnswerB

BigQuery ML supports logistic regression for binary classification directly via SQL. The CREATE MODEL statement with model_type='LOGISTIC_REG' trains the model without Python, and ML.PREDICT generates predictions in SQL. Logistic regression is inherently explainable because coefficients indicate feature importance and direction, satisfying the explainability requirement.

Why this answer

BigQuery ML enables training and prediction using SQL, which perfectly matches the team's need to avoid Python. Logistic regression is a standard binary classification model that provides interpretable coefficients, satisfying explainability. Creating the model with CREATE MODEL and scoring with ML.PREDICT keeps the entire workflow within BigQuery, leveraging its scalability for 2 million rows.

Other options either require coding, use inappropriate algorithms, or involve external services that complicate the low-code requirement.

Exam trap

The trap here is assuming that low-code solutions like AutoML are always the best fit, overlooking that BigQuery ML can handle the entire workflow in SQL without any external tools.

62
MCQeasy

A company needs to extract text from scanned invoices and parse key fields like invoice number and total amount. Which Document AI processor should they use?

A.OCR Processor
B.Contract Parser
C.Form Parser
D.Invoice Parser
AnswerD

Invoice Parser combines OCR with pre-trained extraction of invoice-specific fields such as invoice number and total amount, satisfying the requirement to parse key fields from scanned invoices. Unlike generic OCR, which returns raw text only, it maps recognised content directly to structured invoice schema, removing custom model training.

Why this answer

The Invoice Parser is a specialized Document AI processor pre-trained to extract structured data from invoices, including fields like invoice number, total amount, due date, and line items. It goes beyond generic OCR by understanding the invoice layout and returning key-value pairs. OCR Processor only extracts raw text without parsing fields, while Form Parser and Contract Parser are for general forms and contracts, respectively.

Exam trap

The trap here is confusing generic OCR with specialized parsing; candidates might think OCR is sufficient for extracting fields, but it only provides raw text without structure.

How to eliminate wrong answers

Option A is wrong because the OCR Processor only performs optical character recognition to extract text, but does not parse specific fields like invoice number or total amount. Option B is wrong because the Contract Parser is designed for legal contracts, not invoices, and would not reliably extract invoice-specific fields. Option C is wrong because the Form Parser is a generic form processor that can extract key-value pairs but is not pre-trained specifically for invoices, so it may require custom training and lacks invoice-specific fields.

63
MCQmedium

A company wants to build a recommendation system that suggests products to users based on their past purchase history. They have a large dataset of user-item interactions in BigQuery and want to use a low-code approach. They decide to use BigQuery ML's matrix factorization model. Which SQL statement correctly creates such a model?

A.CREATE MODEL `project.dataset.recommender` OPTIONS(model_type='kmeans', num_clusters=10) AS SELECT user_id, product_id, rating FROM `project.dataset.interactions`
B.CREATE MODEL `project.dataset.recommender` OPTIONS(model_type='logistic_reg', user_col='user_id', item_col='product_id', rating_col='rating') AS SELECT user_id, product_id, rating FROM `project.dataset.interactions`
C.CREATE MODEL `project.dataset.recommender` OPTIONS(model_type='matrix_factorization', user_col='user_id', item_col='product_id', rating_col='rating') AS SELECT user_id, product_id, rating FROM `project.dataset.interactions`
D.CREATE MODEL `project.dataset.recommender` OPTIONS(model_type='matrix_factorization', input_label_cols=['rating']) AS SELECT user_id, product_id, rating FROM `project.dataset.interactions`
AnswerC

This statement correctly creates a matrix factorization model in BigQuery ML. It specifies the model type as matrix_factorization and identifies the user column, item column, and rating column. The SELECT statement provides the training data with user-item interactions and ratings. This is the standard syntax for building a recommendation model using collaborative filtering in BigQuery ML.

Why this answer

The correct SQL statement creates a matrix factorization model in BigQuery ML, specifying the user column, item column, and rating column. This model uses collaborative filtering to learn latent factors for users and items, enabling personalized recommendations. It is a low-code solution that requires only SQL and scales to large datasets, making it ideal for the company's needs.

Exam trap

The trap here is using incorrect options like input_label_cols for matrix factorization, which requires user_col, item_col, and rating_col.

64
MCQeasy

A media company wants to transcribe audio files from customer support calls into text for analysis. The audio is in English with clear speech and no background noise. They want a quick solution with no ML model training. Which Google Cloud service should they use?

A.Translation API to translate the audio
B.AutoML NLP to train a transcription model
C.Vertex AI Workbench to train a custom speech recognition model
D.Speech-to-Text API with the latest_long model
AnswerD

The latest_long model suits extended support-call audio, delivering accurate English transcription without any model training. It satisfies the stem's demand for a quick, pre-trained solution, unlike custom models requiring data and tuning. Speech-to-Text handles clear speech directly, so no ML expertise or training pipeline is needed.

Why this answer

The Speech-to-Text API is Google Cloud's fully managed, pre-trained automatic speech recognition (ASR) service that requires no model training. The latest_long model is specifically optimized for transcribing long-form audio content such as customer support calls, providing high accuracy for clear English speech. Since the audio is clear with no background noise and the company wants a quick, training-free solution, Speech-to-Text API with latest_long is the ideal fit.

Exam trap

PMLE often tests the distinction between pre-trained APIs and custom training services; candidates may incorrectly assume that any ML task requires training a custom model, overlooking that Speech-to-Text API is a ready-to-use pre-trained service.

How to eliminate wrong answers

Option A is wrong because the Translation API translates text between languages and cannot process audio input at all. Option B is wrong because AutoML NLP is designed for text classification, entity extraction, and sentiment analysis—not audio transcription—and it would require labeled training data. Option C is wrong because Vertex AI Workbench is a notebook-based development environment for building and training custom ML models, which contradicts the requirement for no model training and a quick solution.

65
Multi-Selecteasy

A data analyst wants to use BigQuery ML to train a linear regression model (LINEAR_REG) to predict house prices. They have a table with features like square footage, number of bedrooms, and location. Which TWO statements about the training process are correct?

Select 2 answers
A.The analyst must call ML.TRAIN after CREATE MODEL to start training
B.The trained model is stored in Cloud Storage
C.The model must be exported to Vertex AI for prediction
D.The model is automatically evaluated on a held-out test set if data splitting is enabled
E.Training is performed using the CREATE MODEL statement
AnswersD, E

Enabling data splitting in CREATE MODEL reserves a portion of the input rows as a held-out test set, and BigQuery ML automatically computes evaluation metrics on it after training, satisfying the requirement to assess model quality without a separate manual step.

Why this answer

When data splitting is enabled in BigQuery ML, the `CREATE MODEL` statement automatically reserves a portion of the input data as a held-out test set. After training completes, BigQuery ML evaluates the model on this test set and reports metrics like mean absolute error and R², without requiring any manual split or separate evaluation step.

Exam trap

A common misconception is that BigQuery ML requires an explicit training command (like `ML.TRAIN`) or that models are stored in Cloud Storage by default, when in fact training is fully encapsulated in `CREATE MODEL` and models reside in BigQuery's internal storage.

66
MCQeasy

A hospital wants to build a system that automatically transcribes doctors' dictated notes into text and then identifies key medical terms such as diagnoses and medications. They have no ML expertise and want to use Google Cloud's pre-trained APIs. Which combination of services should they use?

A.Speech-to-Text API for transcription and AutoML Natural Language for custom entity extraction.
B.Speech-to-Text API for transcription and Healthcare Natural Language API for medical term extraction.
C.Text-to-Speech API for transcription and Natural Language API for entity extraction.
D.Dialogflow for transcription and Healthcare Natural Language API for medical term extraction.
AnswerB

Speech-to-Text API converts audio to text accurately, and the Healthcare Natural Language API is specifically designed to extract medical entities like diagnoses and medications from text. This combination uses fully managed pre-trained models, requiring no ML expertise, and directly addresses both transcription and medical term identification in a single pipeline.

Why this answer

The hospital needs two capabilities: converting audio to text and extracting medical terms. Speech-to-Text API handles audio transcription with high accuracy, and Healthcare Natural Language API is purpose-built for medical text analysis, identifying entities like medications and diagnoses. Both are pre-trained, requiring no ML expertise, and integrate easily.

Other options use incorrect services for transcription or require custom model training, which is unnecessary given the availability of specialized pre-trained APIs.

Exam trap

The trap here is confusing Text-to-Speech with Speech-to-Text, or assuming that general Natural Language API can handle medical terminology as well as the specialized Healthcare Natural Language API.

67
MCQeasy

A retail company wants to build a demand forecasting model for thousands of product SKUs. They have historical sales data in BigQuery and limited ML expertise. They want to minimize coding and automatically handle seasonality and promotions. Which approach should they use?

A.Use BigQuery ML with the ARIMA_PLUS model type and provide the time series data.
B.Use Dataflow to preprocess the data and then train a custom TensorFlow model on Vertex AI.
C.Use BigQuery ML with the LINEAR_REG model type and include date features.
D.Use Vertex AI AutoML Forecasting with the sales data exported to Cloud Storage.
AnswerA

BigQuery ML's ARIMA_PLUS is designed for univariate time series forecasting and automatically handles seasonality, holiday effects, and outliers. It requires only SQL, making it ideal for teams with limited ML expertise. It also supports large-scale training across many time series, such as product SKUs, and can incorporate covariates like promotions. This directly meets the company's need for minimal coding and automatic seasonality handling.

Why this answer

BigQuery ML's ARIMA_PLUS is purpose-built for time series forecasting and automatically addresses seasonality, holidays, and outliers without manual feature engineering. It operates entirely within BigQuery using SQL, aligning with the team's limited ML expertise and desire to minimize coding. It scales to many time series, making it suitable for thousands of SKUs.

Exam trap

The trap here is assuming that any BigQuery ML model can handle time series seasonality, when only specialized model types like ARIMA_PLUS provide that capability out of the box.

68
MCQmedium

An engineer needs to perform sentiment analysis on customer reviews. They have a large volume of text and need a solution that requires minimal customisation. Which option is most efficient?

A.Use Vertex AI Prediction with a pre-trained model
B.Use BigQuery ML with LOGISTIC_REG
C.Train a custom model using AutoML NLP
D.Use the Natural Language API
AnswerD

The Natural Language API provides pre-trained sentiment analysis, satisfying the minimal-customisation constraint without labelled data or model training. It handles large text volumes through a managed endpoint, unlike custom model approaches that demand dataset preparation and tuning. This makes it the most efficient fit for the stated scenario.

Why this answer

The Natural Language API provides pre-built sentiment analysis with minimal setup. AutoML NLP would require custom training, BigQuery ML is for tabular data, and Vertex AI Prediction needs a deployed model.

Ready to test yourself?

Try a timed practice session using only Pmle Low Code Ml questions.