Courseiva

PDE Ingesting and Processing the Data Practice Question

A media analytics team needs to copy 200 TB of historical log files from an on-premises NFS server into Cloud Storage, then transform them with a Dataflow batch job. The NFS server is reachable only from the corporate network, and the team wants to minimize transfer time and cost. Which approach should they use?

⚠ Common exam trap

The trap here is assuming a VPN plus command-line copy is sufficient for large on-premises transfers, when agent-based Storage Transfer Service is the scalable, managed option.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Storage Transfer Service with an agent pool deployed in the corporate network to transfer directly from the NFS server to Cloud Storage.

Storage Transfer Service with an agent pool is purpose-built for moving large datasets from on-premises sources like NFS to Cloud Storage. Agents run inside the corporate network, so no inbound firewall changes are needed, and the service handles parallelism, retries, and integrity checks. Manual rsync, Transfer Appliance, and single-machine VPN copies are slower or inappropriate at this scale.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create a Cloud VPN tunnel and use gcloud storage cp from an on-premises machine to upload the files.

    Why it's wrong here

    Using gcloud storage cp over a VPN from a single machine does not parallelize across many workers and is limited by the VPN tunnel bandwidth and the source machine's resources. For 200 TB this would take far longer and lacks retry and integrity management, making it unsuitable compared to an agent-based transfer service.

  • ✗

    Mount the NFS share on a Compute Engine VM and run gsutil -m rsync to copy the data to Cloud Storage.

    Why it's wrong here

    Mounting NFS on a VM requires network connectivity and manual retry handling, and gsutil rsync is not optimized for 200 TB with parallel resumable transfers at scale. It also lacks the built-in scheduling, integrity verification, and agent-based architecture that Storage Transfer Service provides for on-premises sources, so it is slower and more error-prone.

  • ✓

    Use Storage Transfer Service with an agent pool deployed in the corporate network to transfer directly from the NFS server to Cloud Storage.

    Why this is correct

    Storage Transfer Service supports on-premises sources through agent pools, which run in the corporate network and pull data from NFS to Cloud Storage. This avoids staging through a VPN and is designed for large-scale, reliable transfers with scheduling and integrity checks, making it the right fit for 200 TB from an NFS source.

  • ✗

    Use Transfer Appliance to ship the data to Google, then load it into Cloud Storage.

    Why it's wrong here

    Transfer Appliance is intended for very large migrations where network transfer is impractical, often hundreds of terabytes to petabytes. For 200 TB on a reachable NFS server, an online agent-based transfer is faster and avoids shipping hardware. Transfer Appliance also does not support incremental or scheduled transfers from a live NFS share.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.