PDE • Practice Test 6
Free PDE practice test — 10 questions with explanations. Set 6. No signup required.
An e-commerce company runs a daily batch pipeline that processes clickstream data from Cloud Storage using Cloud Dataproc with Spark. The pipeline includes a join between a large fact table and a small dimension table. The dimension table is stored in Cloud Storage as a CSV file. The join is slow due to shuffling. The data engineer considers broadcasting the dimension table. However, the dimension table is updated daily and the pipeline reads the latest version. What is the best approach to implement this optimization?
Choose an answer to begin — your selection is scored in the full session.
10 questions · instant feedback and full explanations after every question.