Skip to content

[pull] memtable_as_log_index from topling:memtable_as_log_index - #37

Open
pull[bot] wants to merge 52 commits into
hugegraph:memtable_as_log_indexfrom
topling:memtable_as_log_index
Open

pull[bot] wants to merge 52 commits into
hugegraph:memtable_as_log_indexfrom
topling:memtable_as_log_index

Conversation

@pull

@pull pull Bot commented Sep 13, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

@pull pull Bot locked and limited conversation to collaborators Sep 13, 2026
@pull pull Bot added the ⤵️ pull label Sep 13, 2026
Borrow NeedsUserKeyCompareInGet so Topling pins the writer token
on the caller's hint and skips per-prefix insert_hints_.
fillseq pins one hint through FinishHint for Topling tokens and
concurrent heap splices; SkipList sequential splices are arena-owned.
--use_concurrent_insert selects the Concurrent insert APIs
independently of hint.
The OffsetSkipList fillseq pass stays. CSPP gets the same
workload, sampling, and pages sections so the two memtables
can be compared on equal footing.
Range-del memory is zero when the table is empty; do not touch it on
the flush check. Cover the no-range-del path with a unit test.
Non-concurrent hinted Adds skip the per-key flush check; the inserter
updates flush state once when the hint map is torn down.
The filter is usually absent and first_seqno is set after the first key.
Keep the pages compact by dropping standalone us/op rows and appending
those numbers to the Elapsed time cell.
Use ToplingDB's cheap, millisecond-level ConvertToSST on close when all
MemTables support it, avoiding WAL replay on the next open.

Upstream RocksDB skips flushing WAL-backed MemTables on close because
BuildTable is expensive and can cause a long wait. ToplingDB's MemTables
support much cheaper, millisecond-level ConvertToSST, but the existing
close path does not take advantage of this capability.

Without this change, convertible MemTables are still left for WAL replay
on the next open, adding avoidable recovery work and startup latency
despite the availability of cheap ConvertToSST.
Honor small max_open_files values, including 0, so callers can request a
cheap DB open without unnecessary SST opens. This matters for short-lived
operations that do not read table data.

The existing behavior silently raises this limit to 20 and still opens
SSTs even when opening them for statistics has been disabled. This
defeats the caller's explicit request and violates the principle of
least surprise.

Without this change, callers cannot avoid unnecessary SST opens.
Short-lived DB opens therefore incur avoidable I/O, latency and resource
usage, particularly when accessing SSTs is expensive.
Introduce memtable_crash_safe_recover, disabled by default, so users can
explicitly choose recovery from crash-safe MemTables.

Crash-safe recovery and Flush/Close ConvertToSST serve different purposes.
A separate option keeps recovery opt-in without tying it to
memtable_as_log_index or changing existing WAL-replay behavior by default.
Add methods to MemTableRepFactory for reusing crash-safe file mmap
MemTables during recovery.

The DB recovery path needs these methods to reuse the on-disk files of
MemTable data across different implementations without knowing their
file formats.
Otherwise it must rely on plugin-specific code or rebuild recoverable
data through WAL replay.

Factories that do not support crash-safe recovery remain opted out by
default.
Require compatible WAL settings and WAL-backed writes for MemTable
crash-safe recovery. Disable this recovery mode for read-only Open and
warn that full WAL replay may be very slow.

Recovery reuses on-disk MemTable files and relies on WAL data to recover
writes beyond the published sequence. Disabling WAL, retaining writes
in a process-local WAL buffer, or recycling WAL files would invalidate
those assumptions. WAL compression is unsupported both by crash-safe
recovery and memtable_as_log_index.

Read-only Open cannot run ConvertToSST recovery and install its results.
It must replay WAL data instead, which can take much longer with the large
MemTables encouraged in crash-safe mode. A warning is necessary
even without a configured logger so this potential delay is not silent.
Allow a fresh WAL reader to resume at a saved record boundary while
preserving absolute record offsets for both classic and log-index WALs.

Crash-safe recovery needs to replay the WAL tail without reading the
prefix already covered by recovered MemTable data. Skipping bytes only
in the underlying file loses block alignment and absolute record offsets,
while replaying that prefix defeats the recovery speedup for large
MemTables.
Principle

ToplingDB's file-mmap memtable survives a process crash. This gives the DB an
essential ability to recover.
⇒ 1. But the existing recovery model is inherited from RocksDB, of which the
     approach is replaying the WAL.
⇒ 2. To utilize the crash-safe ability, there is a gap: a crash does not reveal
     how far the memtable insert went. The crash-safe boundary the memtable can
     guarantee is each KV, which does not match the DB's atomic granularity ---
     a WriteBatch.
⇒ 3. So we need to align the memtable's KV boundary to the WriteBatch, this just
     needs a SeqNum.
⇒ 4. The SeqNum must represent a safe point, we call it pubseq, the
     LastPublishedSequence is the ideal candidate for pubseq, our solution is to
     get pubseq as close as possible to LastPublishedSequence ---- this may need
     hiding a few KVs already inserted into the memtable as a sacrifice.
⇒ 5. Once we do not truly reach LastPublishedSequence, what is hidden/sacrificed
     in memtable will be an entire WAL Record (containing at least one
     WriteBatch), then in recovery, we need to replay the tail records in WAL,
     --- an extremely small part of the WAL.

Support

1. On the default write path, an idle DB replays no WAL records. WALs older than
   the KV in the WAL that makes pubseq visible are skipped. The WAL of this KV
   is opened and seeked to the end, and no further record is read.
2. two_write_queues without seq_per_batch publishes after the memtable side
   finishes. An idle DB also replays no records.
3. unordered_write and pipelined_write keep in memory the latest KV in the WAL
   that makes pubseq visible. CSPUBSEQ trails one WAL group, so that last
   group is replayed after a write.
4. allow_2pc still converts leftovers. Records before the KV in the WAL that
   makes pubseq visible are read and not inserted, so recovered_transactions_ is
   rebuilt. 2PC can still be optimized further. This commit does not do that.
5. TransactionDB WriteCommitted is supported, including prepares on the second
   queue. WritePrepared and WriteUnprepared fall back to full WAL replay.

KV in the WAL that makes pubseq visible

1. Publish pubseq and the KV in the WAL that makes it visible, in CSPUBSEQ.
   Convert leftover file-mmap memtables up to pubseq, then replay only the WAL
   tail.
2. Unusable leftovers or an unusable KV in the WAL that makes pubseq visible
   fall back to full WAL replay.
3. An Open whose WAL kind differs from the saved kind of the KV in the WAL that
   makes pubseq visible is not handled here.
When the requested memtable_as_log_index differs from the saved
wal_offset_kind, the previous commit did not handle this scenario.

In this commit:
⇒ 1. If CSPUBSEQ has an established kind and it differs, Open first in the
     saved kind, flush, delete CSPUBSEQ, then Open in the requested kind.
⇒ 2. Prepared transactions recovered on that first Open stop the switch.
     Resolve them, then Open again. With allow_2pc and none left, the
     old-format WALs are retired so the next Open does not replay them into
     the new kind.
⇒ 3. A kind stamped but never used for a WAL is not established, so this path
     does not run. An odd generation does not hide the kind.
⇒ 4. DBImpl::Open still rejects the mismatch. DB::Open and TransactionDB::Open
     run the switch first.
Fill one memtable to 75% of 2GiB, then _exit without Close. Time only
DB::Open on a sparse copy of that directory. Numbers are in
tools/crash_recover_bench.md.

Crash-safe path, /dev/shm, 5 runs. CSPP 9,664,512 keys. Default SkipList
9,090,048 keys. Both configurations use the same ToplingDB build, not
an independently built upstream RocksDB binary.

| recovery path | five runs (ms) | avg (ms) | ratio |
| --- | --- | ---: | ---: |
| CSPP crash-safe | 6.545, 4.692, 3.820, 4.768, 5.120 | 4.989 | 1.0 |
| Default SkipList / WAL replay | 6124.357, 5849.884, 5984.324, 5943.391, 6033.802 | 5987.152 | 1200.1 |
Crash-safe publication relies on writing the 32-byte record containing
pubseq in one instruction. A non-volatile AVX intrinsic does not preserve
that requirement: Clang can split it into two 16-byte stores, leaving a
mixed record if the process crashes between them.

Clang/LLVM specifies that the backend must not split or merge target-legal
volatile loads/stores. AVX provides a native 32-byte store, so volatile
preserves the required instruction width and allows Clang to use the same
publication path as GCC instead of the odd/even fallback.

https://llvm.org/docs/LangRef.html#volatile-memory-accesses
The recovery guide explains database behavior, while the companion article explains why wait-free read structures can remain readable after a process crash. Make both discoverable from the main repository, with English readers directed to the new English editions.
Finding leftover files only tells recovery which files still exist, not
whether any required file has disappeared. Fast recovery needs an
authoritative record of the complete set before it can safely skip WAL
replay.

That record must survive MANIFEST rotation and atomic edits, including the
transition to an empty set. An empty tracked set is not the same as an old
MANIFEST that never tracked MemTables. Required tags make older binaries
reject this state instead of silently ignoring it.

Backing files also occupy DB file numbers and need the same committed
ownership rules as SSTs: retirement must not delete a file that an atomic
conversion has just installed under the same number.
A complete MANIFEST registry is useful only if each MemTable is registered
before it accepts writes. Registration belongs to creation and installation,
not the insert hot path; MemTables prepared ahead of a flush must obey the
same rule before the cache can supply them.

Giving the backing file its final SST number lets ConvertToSST transfer
ownership without a rename gap. The MANIFEST can retire the MemTable and
install its SST together. Until then the mutable file must remain protected
from garbage collection without being exported as an immutable SST.

Recovery can now compare against the registered set and fall back to full
WAL replay when a required file is missing or unusable. Conversion of a
leftover is repeatable, so interruption before the MANIFEST commit must
leave the published recovery cursor usable for another attempt.
Conversion-mode changes must not invalidate the configuration assumptions
made when a factory is created. Both MemTable factories and their table
factory wrappers expose runtime updates, so the same restriction must
hold through both entry points without partially changing other options.

Other runtime tuning and repeated submission of the current mode remain
valid; a dangerous-update opt-in no longer enables mode transitions.
ConvertToSST avoids the cost of rebuilding a table. Falling back to
BuildTable after conversion fails hides the original error and unexpectedly
performs that expensive work. Preserve the conversion failure for the caller
to handle.
MemTable already caches its conversion capability at construction.
Reusing that flag avoids an unnecessary pointer chase on each capability
check.
Unit tests must exercise the same MemTable cache as production.
Compiling that cache out hides allocation, configuration, and
handoff behavior that the production write path depends on.
Failure injection should replace the result of an operation that can
actually fail. A freshly constructed OK status bypasses the seek and
does not exercise the recovery path after a real seek failure.
Conversion mode is fixed when the factory is created. Keeping its
meaning in the common Rep and factory classes gives callers one
consistent capability contract instead of duplicated virtual checks
and independently interpreted state in each implementation.
ConvertToSST converts one existing MemTable rather than rebuilding a
merged table. Waiting for several MemTables is incompatible with that
contract. Normalize the merge threshold for converting factories
while retaining the requested threshold for ordinary MemTables.
FileMmap MemTables already have a backing file for ConvertToSST.
Rebuilding their contents through MemPurge is outside that lifecycle
and must not replace the existing file-backed table.
Crash-safe recovery requires every column family to provide
recoverable FileMmap data. Accepting an incompatible factory would
silently defeat that guarantee. Read-only and secondary opens must
not create writable FileMmap backing files; disabling crash-safe
recovery still leaves ordinary read-write factory choices unrestricted.
Recovery needs a durable inventory to distinguish a missing MemTable
file from one that never existed. These records describe that inventory
without treating mutable files as installed SSTs. Older readers must
reject the records rather than silently ignore required recovery state.
The MANIFEST inventory must survive replay, rollover, and an empty
live set. Failed edits must not change it, and replaying historical
deletions must not schedule live-file cleanup. Keeping the inventory
with its column family also makes file retirement and same-number
SST ownership transfer follow the committed state.
Lock stripe counts are runtime values, so modulo adds integer division
to the transaction lock hot path. It also changes the hash bits used for
stripe selection: with 16 stripes, k1, k2, and k3 all map to the same
stripe, introducing mutex contention between unrelated keys and causing
zero-timeout locking to fail.
AVX publishes the record with one store and does not need the odd/even
protocol used by scalar publication. Sharing that protocol obscures the
different guarantees and adds unnecessary work to the AVX write path.

An interrupted scalar publication can leave an odd generation. Successful
recovery must discard that incomplete cursor and establish an even initial
generation before writes resume; otherwise AVX increments would preserve
the invalid state indefinitely.
@rockeet
rockeet force-pushed the memtable_as_log_index branch from 952c09d to e19e1ab Compare October 5, 2026 14:24
A directory scan cannot reveal a missing member of a MemTable set.
Recoverable files need stable DB identities and a committed inventory
before accepting writes. The DB must retain those files until flush
or recovery commits their retirement, including failed conversions
and another process crash during recovery.

Creation, cache handoff, garbage collection, SST installation, and
recovery share this ownership contract and must become active together.
Normal switches use precreated, registered tables; the rare cache-miss
path may register synchronously instead of waiting indefinitely.
A cached table can leave the queue while its registration is still
in flight. Tests must exercise those interleavings, cache misses, and
registration failures so an uncommitted file is neither exposed for
writes nor removed by garbage collection.
A completed conversion is not yet a committed recovery transition.
Crashes on either side of a MANIFEST commit must preserve a usable
inventory, and an interrupted recovery must remain retryable. Cover
these windows for registration, ordinary flush, atomic multi-CF flush,
and repeated recovery rather than testing only clean reopen.
@rockeet
rockeet force-pushed the memtable_as_log_index branch from e19e1ab to 048c2c9 Compare October 5, 2026 14:42
rockeet added 11 commits October 6, 2026 00:26
A successful table reopen does not detect missing CSPP header metadata:
the table reader can obtain its format information elsewhere. Validate the
serialized header and its checksum so DumpMem cannot silently omit them.
FileDescriptor sequence bounds belong to a DB's version state, not to
every reader of an SST. A recovered table must expose the same published
prefix through DB readers, SstFileReader, and SstFileDumper regardless of
caller-provided bounds.

Ordinary converted tables must remain unfiltered. TableReader views must
also distinguish finite pubseq values from kMaxSequenceNumber.
Keep short statements together without allowing already-wrapped statements
to become unnecessarily wide. Apply this rule only to changed code so
formatting does not obscure the functional diff or disturb existing code.
Recovery statistics must not depend on whether a writer was still alive,
had exited, or had reported its pending WAL bytes before the process died.
Check both MemTable implementations against counts derived from the writes,
including normal conversion and the cached memory estimate.

Header aggregate counters are not the recovery source of truth. Clearing
or inflating them must not affect totals rebuilt from cumulative writers,
and repeating recovery must produce the same table and WAL statistics.
Per-writer statistics must include all successful concurrent insertions
without including a rejected duplicate. Compare recovered table metadata
with the MemTable wrapper and the known workload, covering both ordinary
insertion and insertion hints in CSPP and OffsetSkipList.
Recovery is not an atomic operation and can itself be interrupted. A later
Open must be able to rebuild the same totals even if the mapped aggregates
were left cleared or only partly accumulated.

Exercise interruptions after reset, count, byte and node aggregation with
live and exited writers and with multiple WALs. This protects the separation
between permanent cumulative counters and replaceable recovery output.
Crash-safe recovery can retain physical records beyond pubseq. Recovered statistics may count those records, while iterators correctly hide them. Treating this expected difference as corruption prevents otherwise valid recovered tables from being compacted.

Keep strict record-count validation for exact tables. A recovered reader can declare that num_entries is not exact for its visible contents; failure to obtain that declaration must not suppress a mismatch.
Crash-safe recovery converts the leftovers that a previous avoid-flush
close retained, and retires their inventory entries. A DB configured with
avoid_flush_during_shutdown therefore repeats that cycle on every restart,
and the second cycle depends on recovery leaving a freshly registered
MemTable behind.

Extend the existing avoid-flush case with a second close and reopen. A
recovered MemTable left unregistered would silently degrade that open to a
full WAL replay: no error, no data loss, just the end of crash-safe
recovery.
1. A cspp rep configured with convert_to_sst=kFileMmap is created through
   the overload that takes the memtable file path; the plain overload
   leaves the path empty and cannot build one. Add a flag to select that
   overload, with the path resolved against the factory's chroot_dir.

2. Report ApproximateMemoryUsage after each benchmark, next to the
   arena's, so the mmap-backed working set is visible without guessing
   from bytes per key. The rep is not null there: a benchmark list
   starting with readrandom or readseq dereferences it inside Run() first.
Random key generation can obscure the cost of MemTable lookups, and sampling with replacement does not visit every key. Keep the sampled workload while adding an independently shuffled full-key workload, so lookup cost and full-range coverage can both be measured.

Use big-endian integer keys so bytewise key order matches numeric order.
Constructing LookupKey and a comparator for each lookup adds benchmark overhead that CSPP and OSL do not require. Exercise their native GetPIK path so the measurement better reflects lookup cost, while retaining the comparison-based path for other MemTable representations.
@rockeet
rockeet force-pushed the memtable_as_log_index branch from 05edd78 to 7951012 Compare October 7, 2026 12:53
Forward and reverse traversal have different costs, especially for skip lists.
A separate readreverse workload lets one benchmark invocation measure both
directions against the same populated MemTable, avoiding another fill just
to change the scan direction. Preserve the existing --reverse behavior.
Keep the checked-out sample aligned with its file-mapped MemTable setup so
process restarts can reuse MemTable data through crash-safe recovery.
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant