A data engineer needs to transfer 5 PB of historical data from an on-premises Hadoop cluster to Cloud Storage. The network bandwidth is limited to 1 Gbps, and the transfer must complete within 30 days. Which transfer method should they use?
Transfer Appliance is a physical, shippable device for offline bulk migration. At 5 PB over a 1 Gbps link, online transfer would take far longer than 30 days, so shipping appliances satisfies the bandwidth and deadline constraints.
Why this answer
The Transfer Appliance is a physical device designed for large-scale data transfers when network bandwidth is insufficient. With 5 PB of data and a 1 Gbps link, the theoretical maximum transfer time is over 500 days (5 PB × 8 bits/byte / 1 Gbps / 86400 seconds/day), far exceeding the 30-day window. The Transfer Appliance bypasses network constraints by shipping data physically to Google Cloud.
Option A (gsutil rsync over the internet) is incorrect because it relies on the same 1 Gbps bandwidth, making it impossible to transfer 5 PB within 30 days. Option B (BigQuery Data Transfer Service) is designed for transferring data into BigQuery from cloud sources, not for large on-premises data transfers. Option C (Storage Transfer Service for on-premises) is intended for smaller or incremental transfers over a network, not for a 5 PB initial load within 30 days.
Exam trap
The trap here is that candidates may overestimate network transfer speeds or assume that cloud-native services like Storage Transfer Service can handle any volume, ignoring the fundamental bandwidth math that makes physical shipping the only viable option for 5 PB within 30 days.
How to eliminate wrong answers
Option A is wrong because gsutil rsync over the internet at 1 Gbps would take approximately 500 days to transfer 5 PB, which exceeds the 30-day deadline; it also lacks reliability for such massive transfers over a public network. Option B is wrong because BigQuery Data Transfer Service is designed for scheduled imports from SaaS applications (e.g., Google Ads, Amazon S3) and does not support direct on-premises Hadoop transfers. Option C is wrong because Storage Transfer Service for on-premises requires a network connection (typically via a staging bucket or partner interconnect) and still relies on the same 1 Gbps bandwidth, making it impossible to meet the 30-day requirement.