DUE TO SPAM, SIGN-UP IS DISABLED. Goto Selfserve wiki signup and request an account.
...
Currently, there's no transactional operation for compaction. In the current CompactionScannerFactory, if it fails to flush entry log file, or fails to flush ledgerCache, the "data which is already flushed data" wouldn't be deleted, and it will retry the entry log that is being compacted will be retried again for the next time since the log is still there when compaction fail. This is generating , which would generate duplicated data. And
Moreover, if the data entry log being compacted is has long-lived data and the compaction keeps failing for some reason(e.g. corrupted entry, corrupted index), it would cause the BK disk usage keep growing . Adding transactional operation for compaction would address this issue, for example, if the compaction failed for log1, we should roll back the compaction by deleting the data copied from log1 once we use a separate file for compaction.
...
until the either the entry log can be garbage collected, or disk full.
Proposal
Use a separate log for compaction
In order to address the first issue, we can use a separate log file for compaction and have a separate allocation logic. To allocate a compaction log file, we don't have to choose from the writable ledger directories, which is determined by diskWarnThreshold. In fact, as long as there's a ledger directory that has enough disk space for the next compaction log (we can use log size limit), we should be good to allocate. Because when disks are full, bookie must be running in read only mode, only compaction would write to ledger disks.
Add transactional phases for compaction
Once we separate the log file for compaction, we can achieve a transactional compaction operation. By "transactional", we mean that if anything fail at any phases during compaction, we should be able to roll back the current compaction properly, failed compaction would still be able to retry in the next scan, but rolling back the failed compaction would help us clean up the duplicated data.
Add recovery for compaction
Design