A logistics company wants to optimize their delivery routes using historical GPS data. The data is stored in BigQuery and is updated daily. They need to run a complex machine learning model that requires iterative processing over the entire dataset using Apache Spark. The model training takes several hours and must be run weekly. They want to minimize cost and operational overhead. Which approach should they take?
Dataproc Serverless for Spark can read data directly from BigQuery using the BigQuery connector, eliminating the need to export data. It is a serverless solution, so it automatically provisions and scales resources, and charges only for the duration of the job. This minimizes both cost and operational overhead, as there is no cluster to manage. It supports complex Spark MLlib models and iterative processing, making it ideal for this scenario.
Why this answer
Dataproc Serverless for Spark with the BigQuery connector allows the company to run Spark jobs directly on BigQuery data without exporting it. It is serverless, so it minimizes operational overhead and costs by charging only for job execution. It supports complex Spark MLlib models and iterative processing, making it the best fit for weekly training on historical GPS data.
This approach aligns with the goals of minimizing cost and operational overhead while leveraging Spark's capabilities.
Exam trap
The trap here is assuming that data must be exported from BigQuery to Cloud Storage for Spark processing, overlooking the direct BigQuery connector available in Dataproc Serverless.