Courseiva
← Back to Google Professional Data Engineer questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise Google Professional Data Engineer practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

10
scenario questions
PDE
exam code
Google Cloud
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related PDE topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

In the Vertex AI Pipeline component YAML exhibit, the component is designed to evaluate a model and produce metrics. If the threshold_accuracy is set to 0.85, what is the expected behavior of this component?

Exhibit

Refer to the exhibit.

```
# Vertex AI Pipeline component YAML
name: model-evaluation
inputs:
  model_path:
    type: String
  test_data_path:
    type: String
  threshold_accuracy:
    type: Float
    default: 0.85
outputs:
  evaluation_metrics:
    type: Metrics
implementation:
  container:
    image: gcr.io/my-project/eval:latest
    args: [
      --model_path, {inputValue: model_path},
      --test_data_path, {inputValue: test_data_path},
      --threshold_accuracy, {inputValue: threshold_accuracy},
      --output_path, {outputPath: evaluation_metrics}
    ]
```
Question 2hardmultiple choice
Full question →

Refer to the exhibit. A BigQuery dataset is shared with the group 'analysts@example.com' using the IAM policy shown. A user who is a member of this group reports that they cannot run queries on the dataset, though they can see the tables. What is the most likely reason?

Exhibit

Refer to the exhibit.

```json
{
  "bindings": [
    {
      "role": "roles/bigquery.dataViewer",
      "members": [
        "group:analysts@example.com"
      ]
    }
  ]
}
```
Question 3mediummultiple choice
Full question →

Refer to the exhibit. A Dataflow streaming pipeline subscribes to this Pub/Sub subscription. The pipeline occasionally takes more than 10 seconds to process a message. Which behavior will occur?

Exhibit

Refer to the exhibit.

Cloud Pub/Sub subscription configuration:

{
  "name": "projects/my-project/subscriptions/my-sub",
  "topic": "projects/my-project/topics/my-topic",
  "pushConfig": {},
  "ackDeadlineSeconds": 10,
  "messageRetentionDuration": "86400s",
  "expirationPolicy": {
    "ttl": "604800s"
  },
  "enableMessageOrdering": false,
  "retryPolicy": {
    "minimumBackoff": "10s",
    "maximumBackoff": "600s"
  },
  "deadLetterPolicy": {
    "deadLetterTopic": "projects/my-project/topics/dead-letter-topic",
    "maxDeliveryAttempts": 5
  }
}
Question 4mediummultiple choice
Full question →

Refer to the exhibit. What is the cause of this error?

Network Topology
name=my-endpointmachine-type=n1-standard-2min-replica-count=1max-replica-count=3ERROR: (gcloud.ai.endpoints.create) unrecognized arguments:machine-type
Question 5mediummultiple choice
Full question →

Refer to the exhibit. This log entry was generated by Vertex AI Model Monitoring for a production model. What should the data engineer do to address this issue?

Exhibit

{
  "resource": {"type": "ai_platform_endpoint", "labels": {"endpoint_id": "123"}},
  "severity": "ERROR",
  "jsonPayload": {
    "feature_name": "age",
    "monitoring_type": "prediction_drift",
    "drift_score": 0.85,
    "threshold": 0.7
  }
}
Question 6easymultiple choice
Full question →

Refer to the exhibit. A Cloud Build step fails when pushing a Docker image to Artifact Registry. What is the missing IAM role for the Cloud Build service account?

Exhibit

Refer to the exhibit.
```json
{
  "error": "denied: permission denied for us-central1-docker.pkg.dev/my-project/my-repo/my-model:latest"
}
```
Question 7mediummultiple choice
Full question →

Refer to the exhibit. A team uses this Cloud Build configuration to deploy a service to Cloud Run. The deployment step fails with a 'Permission denied' error. What is the most likely cause?

Network Topology
args: ['run'image'region'platform'Refer to the exhibit.```yamlsteps:- name: 'gcr.io/cloud-builders/docker'args: ['build', '-t', 'gcr.io/$PROJECT_ID/my-image', '.']args: ['push', 'gcr.io/$PROJECT_ID/my-image']- name: 'gcr.io/cloud-builders/gcloud'```
Question 8mediummultiple choice
Full question →

Refer to the exhibit. A Dataflow pipeline is failing intermittently with the shown error. Which step should the team take to ensure data quality and prevent such errors?

Exhibit

Refer to the exhibit.
{
  "insertId": "abc123",
  "jsonPayload": {
    "message": "Error processing element: expected integer at field 'temperature', got string 'hot'",
    "workerId": "worker-5",
    "step": "ParseAndValidate"
  },
  "resource": {
    "type": "dataflow_step",
    "labels": {
      "job_id": "job-1234",
      "step_id": "s2"
    }
  }
}
Question 9mediummultiple choice
Full question →

The exhibit shows a Cloud Logging query result. A data engineer sees this log for a streaming Dataflow job. What is the most likely cause?

Exhibit

Refer to the exhibit.

```
resource.type="dataflow_step"
resource.labels.job_id="2023-01-01_000000-12345678"
"worker pool exhausted"
```
Question 10hardmultiple choice
Full question →

A Dataflow pipeline as described in the exhibit has increasing lag. Which optimization is most likely to reduce the lag?

Exhibit

Refer to the exhibit.

Exhibit:
Pipeline description:
- Source: PubSubIO.read()
- Transform: ParDo(Process)
- Window: Window.into(FixedWindows of 1 minute)
- Transform: GroupByKey
- Sink: Write to BigQuery using StreamingInserts
- Estimated throughput: 10MB/s
- Observed lag: increasing

These PDE practice questions are part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style PDE questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.