Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

A data analyst is tasked with gathering data from a legacy system that only exports CSV files. The files contain headers but no data types. Which tool would best facilitate initial data exploration?

⚠ Common exam trap

The trap is choosing a tool that requires predefined schemas (SQL) or is meant for downstream visualization (Tableau) instead of recognizing pandas as the flexible, schema-inferring exploration tool for raw CSVs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Python pandas

Python pandas is the best tool for initial data exploration of CSV files because its read_csv() function automatically infers data types, handles headers, and provides immediate exploratory methods like .info(), .describe(), and .head(). It requires no schema definition upfront, making it ideal for legacy exports with unknown types.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Hadoop

    Why it's wrong here

    Hadoop stores and processes massive datasets across distributed nodes; it cannot parse a small CSV or infer column types for exploration. It is tempting because Hadoop handles large-scale ingestion, and it would be the right choice when the legacy export is too voluminous for a single machine to process.

  • ✗

    Tableau

    Why it's wrong here

    Tableau is a visualisation tool that infers types on import, but it does not provide the column profiling, null counts and type detection needed before loading. It is tempting for quick charts, yet a spreadsheet or profiling tool suits initial exploration of header-only CSVs.

  • ✗

    SQL database

    Why it's wrong here

    A SQL database requires a predefined schema with declared data types, so header-only CSVs cannot be loaded without manual type specification first. It is tempting as a familiar analysis environment, but it suits structured storage after exploration, not the initial profiling stage.

  • ✓

    Python pandas

    Why this is correct

    Python pandas reads CSV files directly with `read_csv`, inferring column data types automatically despite the header-only source. This satisfies the stem's constraint of untyped legacy exports, enabling immediate exploration through `head()`, `info()` and `describe()` without prior schema definition or manual type assignment.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.