DEA-C01 Data Operations and Support Practice Question
A data engineer is troubleshooting a slow-running Amazon Athena query on a large dataset stored in S3. The query scans many small files. Which TWO actions can improve query performance?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Concatenate small files into larger files
Option C is correct because Athena performance is heavily degraded by many small files, since each file incurs overhead for opening, listing, and reading metadata; concatenating small files into larger files (typically 128 MB or more) reduces this per-file overhead and lets Athena scan data more efficiently. Option D is correct because partitioning the data by a frequently filtered column allows Athena to use partition pruning, so it reads only the relevant S3 prefixes instead of scanning the entire dataset, dramatically reducing the amount of data scanned and query time. Option A is incorrect because adding more small files increases metadata and open/close overhead rather than improving performance, even though parallelism is a factor. Option B is incorrect because S3 server-side encryption is transparent to Athena and disabling it does not affect query performance. Option E is incorrect because converting CSV to JSON does not inherently reduce scanned data or file count and JSON is typically more verbose, so it would not improve performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of files to increase parallelism
Why it's wrong here
Increasing file count worsens the small-file problem: Athena's per-file overhead in listing, opening and parsing metadata dominates, so more files add scheduling cost rather than throughput. It is tempting because parallelism genuinely helps when a few large files are split across workers; compaction into fewer, larger files is the fix here.
- ✗
Disable S3 server-side encryption
Why it's wrong here
Disabling S3 server-side encryption removes a security control without affecting Athena's scan volume or file-listing overhead. It tempts when encryption is assumed to add latency, but Athena performance here is governed by file count, size and columnar format.
- ✓
Concatenate small files into larger files
Why this is correct
Merging many small files reduces per-file overhead: Athena must open, list and read metadata for each object, so fewer larger files cut S3 request costs and task scheduling overhead. This directly addresses the small-file scanning inefficiency described in the stem.
- ✓
Partition the data by a frequently filtered column
Why this is correct
Partitioning by a frequently filtered column enables partition pruning, so Athena reads only the relevant S3 prefixes instead of scanning the whole dataset. This directly reduces the bytes scanned that cause the slow query described in the stem.
- ✗
Convert files from CSV to JSON
Why it's wrong here
JSON is row-oriented and typically larger than CSV, so Athena scans more bytes per query rather than fewer. It tempts as a semi-structured format, but converting to Parquet or ORC with compression and partition pruning is what reduces scanned data.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DEA-C01 question from scratch — 1,321 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DEA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C01 exam.