A DBA is troubleshooting a replica set where the secondary is consistently lagging behind the primary by several hours. The DBA notices that the secondary's optime is far behind the primary's optime, and the secondary's oplog window is only 1 hour. The primary has a heavy write workload. Which action should the DBA take to reduce the replication lag and prevent the secondary from falling off the oplog?
Trap 1: Convert the secondary to an arbiter to reduce its workload.
An arbiter does not store data and cannot replicate or serve reads. Converting a lagging secondary to an arbiter would remove a data-bearing member, reducing redundancy and read capacity. It does not help with replication lag because the arbiter does not apply oplog entries. This would worsen the situation by eliminating a replica.
Trap 2: Increase the write concern on the primary to ensure the secondary…
Increasing write concern to require acknowledgment from the secondary would slow down the primary's write throughput and could increase lag if the secondary is already struggling. Write concern controls durability guarantees, not replication speed. It does not help the secondary apply oplog entries faster and may exacerbate the problem by adding latency.
Trap 3: Increase the size of the oplog on the primary and secondaries.
Increasing the oplog size extends the oplog window, which helps prevent the secondary from falling off the oplog and requiring a full resync. However, it does not address the root cause of the lag, which is likely insufficient write capacity or network throughput on the secondary. A larger oplog is a mitigation, not a solution to reduce lag.
- A
Scale the secondary vertically by adding more CPU, RAM, or faster disk I/O.
Replication lag often occurs when the secondary cannot apply oplog entries as fast as the primary generates them. Upgrading the secondary's hardware—especially faster disks and more CPU—can increase its oplog application rate, reducing lag. This addresses the underlying performance bottleneck and is the most direct way to improve replication throughput.
- B
Convert the secondary to an arbiter to reduce its workload.
Why it fails: An arbiter does not store data and cannot replicate or serve reads. Converting a lagging secondary to an arbiter would remove a data-bearing member, reducing redundancy and read capacity. It does not help with replication lag because the arbiter does not apply oplog entries. This would worsen the situation by eliminating a replica.
- C
Increase the write concern on the primary to ensure the secondary acknowledges writes.
Why it fails: Increasing write concern to require acknowledgment from the secondary would slow down the primary's write throughput and could increase lag if the secondary is already struggling. Write concern controls durability guarantees, not replication speed. It does not help the secondary apply oplog entries faster and may exacerbate the problem by adding latency.
- D
Increase the size of the oplog on the primary and secondaries.
Why it fails: Increasing the oplog size extends the oplog window, which helps prevent the secondary from falling off the oplog and requiring a full resync. However, it does not address the root cause of the lag, which is likely insufficient write capacity or network throughput on the secondary. A larger oplog is a mitigation, not a solution to reduce lag.