A data engineer needs to load data from CSV files in Cloud Storage into BigQuery. The CSV files have a header row and some columns contain nested JSON strings. Which TWO methods can they use to load this data into BigQuery?
The Storage Write API can be used to stream data from CSV after parsing.
Why this answer
Option D is correct because a BigQuery load job natively supports the CSV format, including a header row via the --skip_leading_rows parameter, and can ingest Cloud Storage files directly into a native BigQuery table. Option B is correct because the Storage Write API lets a custom application parse the CSV and nested JSON strings itself and then stream the resulting rows into BigQuery, giving full control over transformation during load. Option A is wrong because Datastream is a change data capture (CDC) and replication service for databases such as MySQL, PostgreSQL, Oracle, and SQL Server, not a CSV file loader.
Option C is wrong because an external table (federated query) only queries data in place in Cloud Storage; it does not load the data into BigQuery. Option E is wrong because gsutil only copies objects between Cloud Storage locations and cannot write data into BigQuery tables.
Exam trap
Google often tests the distinction between loading data into BigQuery (permanent storage) versus querying external data sources (federated queries), causing candidates to mistakenly choose Option C as a valid loading method.