AI0-001 AI Models and Data Engineering Practice Question
A data pipeline processes customer data from multiple sources. The data quality check reveals duplicate records. Which step should the pipeline include to handle this?
⚠ Common exam trap
This question tests the distinction between data quality actions (deduplication) and data security or formatting actions (encryption, transformation), leading candidates to confuse validation (which only flags issues) with remediation (which removes duplicates).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Data deduplication
Duplicate records in a data pipeline compromise data integrity and downstream analytics. Data deduplication (Option A) is the correct step because it identifies and removes redundant entries based on key fields or fuzzy matching, ensuring each customer record is unique. This is a core data quality operation in ETL pipelines, often implemented via hash-based comparison or SQL window functions like ROW_NUMBER().
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Data deduplication
Why this is correct
Data deduplication removes duplicate records by identifying matching rows across sources and retaining a single canonical entry, directly resolving the duplicate records the quality check flagged. It satisfies the stem's constraint of handling duplicates within the pipeline, unlike profiling or validation, which detect issues but do not eliminate them.
- ✗
Data encryption
Why it's wrong here
Data encryption protects data at rest or in transit, so it cannot detect or remove duplicate rows; the pipeline still emits repeated records. It is tempting because encryption is a core security control, and it would be the right choice when the requirement is protecting sensitive customer data from unauthorised access, not resolving duplication.
- ✗
Data transformation
Why it's wrong here
Data transformation alters formats, scales or encodes values; it does not identify or remove duplicate rows. Deduplication requires a distinct operation comparing records on key fields. Transformation would be the right step for standardising date formats or unit conversions before loading, which is a separate concern from the duplicate-record requirement here.
- ✗
Data validation
Why it's wrong here
Data validation checks records against rules such as type, range or format, but it does not compare rows against each other, so duplicates pass through untouched. It is tempting because validation is a standard pipeline quality gate, and it would be the right choice for rejecting malformed or out-of-range values rather than repeated records.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.