20+ practice questions focused on Ingesting and Processing the Data — one of the most tested topics on the Google Professional Data Engineer exam. Each question includes a detailed explanation so you learn why the right answer is correct.
Start Ingesting and Processing the Data PracticeA data engineer needs to load 10 TB of CSV files from Amazon S3 into Google BigQuery on a daily basis. Which service should they use to automate this transfer?
Explanation: BigQuery Data Transfer Service is designed to automate data ingestion from external sources, including Amazon S3, into BigQuery on a scheduled basis. It supports CSV and other formats, handles large volumes (10 TB daily), and provides built-in scheduling, monitoring, and retry logic. This makes it the most efficient and managed solution for recurring S3-to-BigQuery loads.
Your company uses Kafka for event streaming. You want to run Kafka on Google Cloud with the ability to auto-scale clusters and use managed infrastructure. Which service should you choose?
Explanation: Dataproc is the correct choice because it is a managed Spark and Hadoop service on Google Cloud that supports running Kafka clusters via initialization actions. It allows auto-scaling of worker nodes and integrates with GCP storage and networking, providing the managed infrastructure required for Kafka event streaming. Note that Confluent Cloud is a third-party managed Kafka service, not a GCP-native service, and Cloud Pub/Sub is a messaging service, not a Kafka replacement. Cloud Dataflow is for data processing pipelines, not for running Kafka itself.
You need to perform a one-time migration of historical data from an on-premises Teradata data warehouse to BigQuery. The data volume is 50 TB and you have a high-speed network connection (10 Gbps). What is the most efficient way to load the data?
Explanation: BigQuery Data Transfer Service for Teradata is designed for this purpose; it can directly connect to Teradata and transfer data to BigQuery. Exporting to CSV then loading via gsutil is possible but less efficient. Transfer Appliance is for offline transfer but you have high-speed network. Dataproc is not needed.
You have a Dataflow pipeline that processes streaming data with high throughput. You notice that the pipeline is experiencing high latency and the workers are underutilized. Which Dataflow feature can automatically optimize resource allocation?
Explanation: Dataflow Prime is Google Cloud's next-generation Dataflow runner that automatically optimizes resource allocation, including right-sizing worker resources and using features like 'Vertical Autoscaling' to adjust memory per worker. It addresses exactly the scenario described — high latency with underutilized workers — by dynamically tuning resources rather than relying on static worker configurations.
Your organization uses dbt (data build tool) for transformations on BigQuery. You need to run dbt models on a schedule and manage versions. Which Google Cloud service can execute dbt jobs in a serverless manner?
Explanation: Cloud Composer is the correct choice because it is a fully managed, serverless (managed) Apache Airflow service on Google Cloud that can schedule and orchestrate dbt jobs, including version management through Airflow's DAGs and connections to BigQuery. It natively supports running dbt commands via BashOperator or KubernetesPodOperator, making it suitable for scheduled dbt transformations on BigQuery. Dataflow is a managed Apache Beam service for stream and batch data processing, not for orchestrating dbt jobs. Cloud Build is a CI/CD service for building and deploying artifacts, not a scheduler for recurring dbt runs. Cloud Run runs containerized applications serverlessly but lacks built-in scheduling and orchestration for dbt workflows.
+15 more Ingesting and Processing the Data questions available
Practice all Ingesting and Processing the Data questions1. Baseline your knowledge
Start with 10 questions to gauge your current understanding of Ingesting and Processing the Data. This tells you whether you need a concept refresher or just practice.
2. Review every explanation
For each question — right or wrong — read the full explanation. Understanding why an answer is correct is more valuable than knowing the answer itself.
3. Focus on exam traps
Ingesting and Processing the Data questions on the PDE frequently use trap wording. Look for subtle differences in answers that test your precision, not just general knowledge.
4. Reach 80% consistently
Do repeated sessions until you score 80%+ three times in a row. Then move to mixed-mode practice to test cross-topic recall under realistic conditions.
The exact number varies per candidate. Ingesting and Processing the Data is tested as part of the Google Professional Data Engineer blueprint. Practicing with targeted Ingesting and Processing the Data questions ensures you can handle any format or difficulty that appears.
Yes. Courseiva provides free PDE practice questions across all exam topics and domains. The platform includes topic-based practice, mock exams, missed-question review, bookmarked questions, and readiness tracking — no account required.
Difficulty is subjective, but Ingesting and Processing the Data is a high-priority exam concept tested in multiple ways — direct recall, scenario analysis, and command-output interpretation. Consistent practice is the best way to build confidence.
Launch a full Ingesting and Processing the Data practice session with instant scoring and detailed explanations.
Start Ingesting and Processing the Data Practice →