DOP-C02 SDLC Automation Practice Question
A company uses AWS CodeCommit for source control. Developers frequently push large binary files (e.g., compiled binaries, datasets) to the repository, causing repository size to grow and clone operations to become slow. What is the BEST approach to manage this?
⚠ Common exam trap
A common mix-up: candidates assume increasing quotas or splitting repositories will solve performance issues, but they fail to recognize that Git LFS is the only option that directly addresses the root cause—large binary files bloating the repository and slowing Git operations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Git LFS in CodeCommit and configure the large files to use LFS.
AWS CodeCommit supports Git Large File Storage (LFS), which replaces large files in the repository with text pointers while storing the actual binary content in a separate hosted storage backend. This keeps the repository lightweight, speeds up clone and fetch operations, and avoids hitting the default repository size limits. Enabling Git LFS is the recommended and best practice for managing large binary files in CodeCommit.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use S3 as the source in CodePipeline and skip CodeCommit for binaries.
Why it's wrong here
Using S3 as the source in CodePipeline would decouple binary assets from the Git history, requiring a custom synchronization mechanism and breaking the atomic relationship between a code commit and the exact binary version it references. CodePipeline treats S3 and CodeCommit as separate source actions, so you lose native versioning, code review, and rollout reproducibility. Developers would also need a parallel workflow to upload and download binaries, which doesn't reduce repository size and makes builds less traceable.
- ✗
Store binaries in a separate CodeCommit repository.
Why it's wrong here
Splitting binaries into a separate CodeCommit repository still requires developers to clone that additional repository, and that repository will grow just as large and slow as the original. Git has no built-in mechanism to reference a specific binary commit from a code commit, so you would need manual or scripted updates to track versions. This approach also multiplies permission management and breaks atomic deployments because the binary repo and the code repo can be out of sync.
- ✗
Increase the repository size limit by requesting a quota increase.
Why it's wrong here
Requesting a repository size increase does not solve the underlying performance problem: Git stores every historical version of every file, so a repo with many large binaries forces every clone, fetch, and push to transfer these large blobs, making operations slow and expensive. CodeCommit's service quota is a hard ceiling, and even if you could raise it, Git clients and CodePipeline have their own practical limits on pack size and memory. Large binary files also degrade local operations like checkout and diff, so increasing the cap only delays the inevitable failure.
- ✓
Enable Git LFS in CodeCommit and configure the large files to use LFS.
Why this is correct
Git LFS in CodeCommit replaces large binary files with small pointer files in the repository and stores the actual binary content in an S3 bucket managed by the service. When developers clone the repo, they only fetch pointers, keeping the repository small and clones fast; the actual binary is downloaded on demand when checking out a specific revision. This maintains Git's full versioning, branching, and commit atomicity while ensuring that the repository itself never bloats.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every DOP-C02 question from scratch — 1,298 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DOP-C02 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DOP-C02 exam.