Repository navigation
Conversation
Borrow NeedsUserKeyCompareInGet so Topling pins the writer token on the caller's hint and skips per-prefix insert_hints_.
fillseq pins one hint through FinishHint for Topling tokens and concurrent heap splices; SkipList sequential splices are arena-owned. --use_concurrent_insert selects the Concurrent insert APIs independently of hint.
The OffsetSkipList fillseq pass stays. CSPP gets the same workload, sampling, and pages sections so the two memtables can be compared on equal footing.
Range-del memory is zero when the table is empty; do not touch it on the flush check. Cover the no-range-del path with a unit test.
Non-concurrent hinted Adds skip the per-key flush check; the inserter updates flush state once when the hint map is torn down.
The filter is usually absent and first_seqno is set after the first key.
Keep the pages compact by dropping standalone us/op rows and appending those numbers to the Elapsed time cell.
Use ToplingDB's cheap, millisecond-level ConvertToSST on close when all MemTables support it, avoiding WAL replay on the next open. Upstream RocksDB skips flushing WAL-backed MemTables on close because BuildTable is expensive and can cause a long wait. ToplingDB's MemTables support much cheaper, millisecond-level ConvertToSST, but the existing close path does not take advantage of this capability. Without this change, convertible MemTables are still left for WAL replay on the next open, adding avoidable recovery work and startup latency despite the availability of cheap ConvertToSST.
Honor small max_open_files values, including 0, so callers can request a cheap DB open without unnecessary SST opens. This matters for short-lived operations that do not read table data. The existing behavior silently raises this limit to 20 and still opens SSTs even when opening them for statistics has been disabled. This defeats the caller's explicit request and violates the principle of least surprise. Without this change, callers cannot avoid unnecessary SST opens. Short-lived DB opens therefore incur avoidable I/O, latency and resource usage, particularly when accessing SSTs is expensive.
Introduce memtable_crash_safe_recover, disabled by default, so users can explicitly choose recovery from crash-safe MemTables. Crash-safe recovery and Flush/Close ConvertToSST serve different purposes. A separate option keeps recovery opt-in without tying it to memtable_as_log_index or changing existing WAL-replay behavior by default.
Add methods to MemTableRepFactory for reusing crash-safe file mmap MemTables during recovery. The DB recovery path needs these methods to reuse the on-disk files of MemTable data across different implementations without knowing their file formats. Otherwise it must rely on plugin-specific code or rebuild recoverable data through WAL replay. Factories that do not support crash-safe recovery remain opted out by default.
Require compatible WAL settings and WAL-backed writes for MemTable crash-safe recovery. Disable this recovery mode for read-only Open and warn that full WAL replay may be very slow. Recovery reuses on-disk MemTable files and relies on WAL data to recover writes beyond the published sequence. Disabling WAL, retaining writes in a process-local WAL buffer, or recycling WAL files would invalidate those assumptions. WAL compression is unsupported both by crash-safe recovery and memtable_as_log_index. Read-only Open cannot run ConvertToSST recovery and install its results. It must replay WAL data instead, which can take much longer with the large MemTables encouraged in crash-safe mode. A warning is necessary even without a configured logger so this potential delay is not silent.
Allow a fresh WAL reader to resume at a saved record boundary while preserving absolute record offsets for both classic and log-index WALs. Crash-safe recovery needs to replay the WAL tail without reading the prefix already covered by recovered MemTable data. Skipping bytes only in the underlying file loses block alignment and absolute record offsets, while replaying that prefix defeats the recovery speedup for large MemTables.
Principle
ToplingDB's file-mmap memtable survives a process crash. This gives the DB an
essential ability to recover.
⇒ 1. But the existing recovery model is inherited from RocksDB, of which the
approach is replaying the WAL.
⇒ 2. To utilize the crash-safe ability, there is a gap: a crash does not reveal
how far the memtable insert went. The crash-safe boundary the memtable can
guarantee is each KV, which does not match the DB's atomic granularity ---
a WriteBatch.
⇒ 3. So we need to align the memtable's KV boundary to the WriteBatch, this just
needs a SeqNum.
⇒ 4. The SeqNum must represent a safe point, we call it pubseq, the
LastPublishedSequence is the ideal candidate for pubseq, our solution is to
get pubseq as close as possible to LastPublishedSequence ---- this may need
hiding a few KVs already inserted into the memtable as a sacrifice.
⇒ 5. Once we do not truly reach LastPublishedSequence, what is hidden/sacrificed
in memtable will be an entire WAL Record (containing at least one
WriteBatch), then in recovery, we need to replay the tail records in WAL,
--- an extremely small part of the WAL.
Support
1. On the default write path, an idle DB replays no WAL records. WALs older than
the KV in the WAL that makes pubseq visible are skipped. The WAL of this KV
is opened and seeked to the end, and no further record is read.
2. two_write_queues without seq_per_batch publishes after the memtable side
finishes. An idle DB also replays no records.
3. unordered_write and pipelined_write keep in memory the latest KV in the WAL
that makes pubseq visible. CSPUBSEQ trails one WAL group, so that last
group is replayed after a write.
4. allow_2pc still converts leftovers. Records before the KV in the WAL that
makes pubseq visible are read and not inserted, so recovered_transactions_ is
rebuilt. 2PC can still be optimized further. This commit does not do that.
5. TransactionDB WriteCommitted is supported, including prepares on the second
queue. WritePrepared and WriteUnprepared fall back to full WAL replay.
KV in the WAL that makes pubseq visible
1. Publish pubseq and the KV in the WAL that makes it visible, in CSPUBSEQ.
Convert leftover file-mmap memtables up to pubseq, then replay only the WAL
tail.
2. Unusable leftovers or an unusable KV in the WAL that makes pubseq visible
fall back to full WAL replay.
3. An Open whose WAL kind differs from the saved kind of the KV in the WAL that
makes pubseq visible is not handled here.
When the requested memtable_as_log_index differs from the saved
wal_offset_kind, the previous commit did not handle this scenario.
In this commit:
⇒ 1. If CSPUBSEQ has an established kind and it differs, Open first in the
saved kind, flush, delete CSPUBSEQ, then Open in the requested kind.
⇒ 2. Prepared transactions recovered on that first Open stop the switch.
Resolve them, then Open again. With allow_2pc and none left, the
old-format WALs are retired so the next Open does not replay them into
the new kind.
⇒ 3. A kind stamped but never used for a WAL is not established, so this path
does not run. An odd generation does not hide the kind.
⇒ 4. DBImpl::Open still rejects the mismatch. DB::Open and TransactionDB::Open
run the switch first.
Fill one memtable to 75% of 2GiB, then _exit without Close. Time only DB::Open on a sparse copy of that directory. Numbers are in tools/crash_recover_bench.md. Crash-safe path, /dev/shm, 5 runs. CSPP 9,664,512 keys. Default SkipList 9,090,048 keys. Both configurations use the same ToplingDB build, not an independently built upstream RocksDB binary. | recovery path | five runs (ms) | avg (ms) | ratio | | --- | --- | ---: | ---: | | CSPP crash-safe | 6.545, 4.692, 3.820, 4.768, 5.120 | 4.989 | 1.0 | | Default SkipList / WAL replay | 6124.357, 5849.884, 5984.324, 5943.391, 6033.802 | 5987.152 | 1200.1 |
Crash-safe publication relies on writing the 32-byte record containing pubseq in one instruction. A non-volatile AVX intrinsic does not preserve that requirement: Clang can split it into two 16-byte stores, leaving a mixed record if the process crashes between them. Clang/LLVM specifies that the backend must not split or merge target-legal volatile loads/stores. AVX provides a native 32-byte store, so volatile preserves the required instruction width and allows Clang to use the same publication path as GCC instead of the odd/even fallback. https://llvm.org/docs/LangRef.html#volatile-memory-accesses
The recovery guide explains database behavior, while the companion article explains why wait-free read structures can remain readable after a process crash. Make both discoverable from the main repository, with English readers directed to the new English editions.
Finding leftover files only tells recovery which files still exist, not whether any required file has disappeared. Fast recovery needs an authoritative record of the complete set before it can safely skip WAL replay. That record must survive MANIFEST rotation and atomic edits, including the transition to an empty set. An empty tracked set is not the same as an old MANIFEST that never tracked MemTables. Required tags make older binaries reject this state instead of silently ignoring it. Backing files also occupy DB file numbers and need the same committed ownership rules as SSTs: retirement must not delete a file that an atomic conversion has just installed under the same number.
A complete MANIFEST registry is useful only if each MemTable is registered before it accepts writes. Registration belongs to creation and installation, not the insert hot path; MemTables prepared ahead of a flush must obey the same rule before the cache can supply them. Giving the backing file its final SST number lets ConvertToSST transfer ownership without a rename gap. The MANIFEST can retire the MemTable and install its SST together. Until then the mutable file must remain protected from garbage collection without being exported as an immutable SST. Recovery can now compare against the registered set and fall back to full WAL replay when a required file is missing or unusable. Conversion of a leftover is repeatable, so interruption before the MANIFEST commit must leave the published recovery cursor usable for another attempt.
This reverts commit cd79ca7.
This reverts commit e1e7809.
Conversion-mode changes must not invalidate the configuration assumptions made when a factory is created. Both MemTable factories and their table factory wrappers expose runtime updates, so the same restriction must hold through both entry points without partially changing other options. Other runtime tuning and repeated submission of the current mode remain valid; a dangerous-update opt-in no longer enables mode transitions.
ConvertToSST avoids the cost of rebuilding a table. Falling back to BuildTable after conversion fails hides the original error and unexpectedly performs that expensive work. Preserve the conversion failure for the caller to handle.
MemTable already caches its conversion capability at construction. Reusing that flag avoids an unnecessary pointer chase on each capability check.
Unit tests must exercise the same MemTable cache as production. Compiling that cache out hides allocation, configuration, and handoff behavior that the production write path depends on.
Failure injection should replace the result of an operation that can actually fail. A freshly constructed OK status bypasses the seek and does not exercise the recovery path after a real seek failure.
Conversion mode is fixed when the factory is created. Keeping its meaning in the common Rep and factory classes gives callers one consistent capability contract instead of duplicated virtual checks and independently interpreted state in each implementation.
ConvertToSST converts one existing MemTable rather than rebuilding a merged table. Waiting for several MemTables is incompatible with that contract. Normalize the merge threshold for converting factories while retaining the requested threshold for ordinary MemTables.
FileMmap MemTables already have a backing file for ConvertToSST. Rebuilding their contents through MemPurge is outside that lifecycle and must not replace the existing file-backed table.
Crash-safe recovery requires every column family to provide recoverable FileMmap data. Accepting an incompatible factory would silently defeat that guarantee. Read-only and secondary opens must not create writable FileMmap backing files; disabling crash-safe recovery still leaves ordinary read-write factory choices unrestricted.
Recovery needs a durable inventory to distinguish a missing MemTable file from one that never existed. These records describe that inventory without treating mutable files as installed SSTs. Older readers must reject the records rather than silently ignore required recovery state.
The MANIFEST inventory must survive replay, rollover, and an empty live set. Failed edits must not change it, and replaying historical deletions must not schedule live-file cleanup. Keeping the inventory with its column family also makes file retirement and same-number SST ownership transfer follow the committed state.
Lock stripe counts are runtime values, so modulo adds integer division to the transaction lock hot path. It also changes the hash bits used for stripe selection: with 16 stripes, k1, k2, and k3 all map to the same stripe, introducing mutex contention between unrelated keys and causing zero-timeout locking to fail.
AVX publishes the record with one store and does not need the odd/even protocol used by scalar publication. Sharing that protocol obscures the different guarantees and adds unnecessary work to the AVX write path. An interrupted scalar publication can leave an odd generation. Successful recovery must discard that incomplete cursor and establish an even initial generation before writes resume; otherwise AVX increments would preserve the invalid state indefinitely.
rockeet
force-pushed
the
memtable_as_log_index
branch
from
October 5, 2026 14:24
952c09d to
e19e1ab
Compare
A directory scan cannot reveal a missing member of a MemTable set. Recoverable files need stable DB identities and a committed inventory before accepting writes. The DB must retain those files until flush or recovery commits their retirement, including failed conversions and another process crash during recovery. Creation, cache handoff, garbage collection, SST installation, and recovery share this ownership contract and must become active together. Normal switches use precreated, registered tables; the rare cache-miss path may register synchronously instead of waiting indefinitely.
A cached table can leave the queue while its registration is still in flight. Tests must exercise those interleavings, cache misses, and registration failures so an uncommitted file is neither exposed for writes nor removed by garbage collection.
A completed conversion is not yet a committed recovery transition. Crashes on either side of a MANIFEST commit must preserve a usable inventory, and an interrupted recovery must remain retryable. Cover these windows for registration, ordinary flush, atomic multi-CF flush, and repeated recovery rather than testing only clean reopen.
rockeet
force-pushed
the
memtable_as_log_index
branch
from
October 5, 2026 14:42
e19e1ab to
048c2c9
Compare
A successful table reopen does not detect missing CSPP header metadata: the table reader can obtain its format information elsewhere. Validate the serialized header and its checksum so DumpMem cannot silently omit them.
FileDescriptor sequence bounds belong to a DB's version state, not to every reader of an SST. A recovered table must expose the same published prefix through DB readers, SstFileReader, and SstFileDumper regardless of caller-provided bounds. Ordinary converted tables must remain unfiltered. TableReader views must also distinguish finite pubseq values from kMaxSequenceNumber.
Keep short statements together without allowing already-wrapped statements to become unnecessarily wide. Apply this rule only to changed code so formatting does not obscure the functional diff or disturb existing code.
Recovery statistics must not depend on whether a writer was still alive, had exited, or had reported its pending WAL bytes before the process died. Check both MemTable implementations against counts derived from the writes, including normal conversion and the cached memory estimate. Header aggregate counters are not the recovery source of truth. Clearing or inflating them must not affect totals rebuilt from cumulative writers, and repeating recovery must produce the same table and WAL statistics.
Per-writer statistics must include all successful concurrent insertions without including a rejected duplicate. Compare recovered table metadata with the MemTable wrapper and the known workload, covering both ordinary insertion and insertion hints in CSPP and OffsetSkipList.
Recovery is not an atomic operation and can itself be interrupted. A later Open must be able to rebuild the same totals even if the mapped aggregates were left cleared or only partly accumulated. Exercise interruptions after reset, count, byte and node aggregation with live and exited writers and with multiple WALs. This protects the separation between permanent cumulative counters and replaceable recovery output.
Crash-safe recovery can retain physical records beyond pubseq. Recovered statistics may count those records, while iterators correctly hide them. Treating this expected difference as corruption prevents otherwise valid recovered tables from being compacted. Keep strict record-count validation for exact tables. A recovered reader can declare that num_entries is not exact for its visible contents; failure to obtain that declaration must not suppress a mismatch.
Crash-safe recovery converts the leftovers that a previous avoid-flush close retained, and retires their inventory entries. A DB configured with avoid_flush_during_shutdown therefore repeats that cycle on every restart, and the second cycle depends on recovery leaving a freshly registered MemTable behind. Extend the existing avoid-flush case with a second close and reopen. A recovered MemTable left unregistered would silently degrade that open to a full WAL replay: no error, no data loss, just the end of crash-safe recovery.
1. A cspp rep configured with convert_to_sst=kFileMmap is created through the overload that takes the memtable file path; the plain overload leaves the path empty and cannot build one. Add a flag to select that overload, with the path resolved against the factory's chroot_dir. 2. Report ApproximateMemoryUsage after each benchmark, next to the arena's, so the mmap-backed working set is visible without guessing from bytes per key. The rep is not null there: a benchmark list starting with readrandom or readseq dereferences it inside Run() first.
Random key generation can obscure the cost of MemTable lookups, and sampling with replacement does not visit every key. Keep the sampled workload while adding an independently shuffled full-key workload, so lookup cost and full-range coverage can both be measured. Use big-endian integer keys so bytewise key order matches numeric order.
Constructing LookupKey and a comparator for each lookup adds benchmark overhead that CSPP and OSL do not require. Exercise their native GetPIK path so the measurement better reflects lookup cost, while retaining the comparison-based path for other MemTable representations.
rockeet
force-pushed
the
memtable_as_log_index
branch
from
October 7, 2026 12:53
05edd78 to
7951012
Compare
Forward and reverse traversal have different costs, especially for skip lists. A separate readreverse workload lets one benchmark invocation measure both directions against the same populated MemTable, avoiding another fill just to change the scan direction. Preserve the existing --reverse behavior.
Keep the checked-out sample aligned with its file-mapped MemTable setup so process restarts can reuse MemTable data through crash-safe recovery.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )