Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions cfg/xrpld-example.cfg
Original file line number Diff line number Diff line change
Expand Up @@ -1085,6 +1085,18 @@
# greater than the current "can_delete" setting.
# Default is 0.
#
# online_delete_generations
# Number of NodeStore generations to keep in the
# rotation ring. Rotation writes to the newest
# generation and reads walk newest to oldest; when
# the ring is full, the oldest generation is retired
# by copying forward only the nodes it still serves,
# then dropping the file. A larger value re-stores a
# cold node less often (once per this many
# rotations) at the cost of keeping more history on
# disk. Clamped to the range 2 to 64.
# Default is 8.
#
# delete_batch When automatically purging, SQLite database
# records are deleted in batches. This value
# controls the maximum size of each batch. Larger
Expand Down
1 change: 1 addition & 0 deletions include/xrpl/config/Constants.h
Original file line number Diff line number Diff line change
Expand Up @@ -135,6 +135,7 @@ struct Keys
static constexpr auto kNormalConsensusIncreasePercent = "normal_consensus_increase_percent";
static constexpr auto kNudbBlockSize = "nudb_block_size";
static constexpr auto kOnlineDelete = "online_delete";
static constexpr auto kOnlineDeleteGenerations = "online_delete_generations";
static constexpr auto kOpenFiles = "open_files";
static constexpr auto kOptions = "options";
static constexpr auto kOverlay = "overlay";
Expand Down
72 changes: 51 additions & 21 deletions include/xrpl/nodestore/DatabaseRotating.h
Original file line number Diff line number Diff line change
Expand Up @@ -5,20 +5,30 @@
#include <xrpl/nodestore/Database.h>
#include <xrpl/nodestore/Scheduler.h>

#include <cstddef>
#include <cstdint>
#include <functional>
#include <memory>
#include <string>
#include <vector>

namespace xrpl::node_store {

/* This class has two key-value store Backend objects for persisting SHAMap
* records. This facilitates online deletion of data. New backends are
* rotated in. Old ones are rotated out and deleted.
/* This class keeps a ring of append-only key-value Backend objects (generations)
* for persisting SHAMap records, to facilitate online deletion of data. New nodes
* are written to the newest (writable) generation; reads probe newest -> oldest.
* Rather than copying the entire live state into a fresh backend every rotation
* (O(total state)), a generation is dropped only once its still-live nodes have been
* evacuated forward, so a cold node is re-stored ~once per ring cycle instead of every
* rotation (O(churn)).
*/

class DatabaseRotating : public Database
{
public:
// Receives the full generation ring, ordered oldest -> newest, to persist durably.
using RingPersist = std::function<void(std::vector<std::string> const& generations)>;

DatabaseRotating(
Scheduler& scheduler,
int readThreads,
Expand All @@ -29,31 +39,51 @@ class DatabaseRotating : public Database
}

/**
* Rotates the backends.
* Append a fresh writable generation. The prior writable becomes a sealed,
* read-only generation that remains in the ring (still served by reads).
*
* @param newBackend New writable backend
* @param f A function executed after the rotation outside of lock. The
* values passed to f will be the new backend database names _after_
* rotation.
* @param newWritable The new (empty) writable backend.
* @param persist Executed after the push, outside the lock, with the full ring
* (oldest -> newest) so the caller can durably record it.
*/
virtual void
rotate(
std::unique_ptr<node_store::Backend>&& newBackend,
std::function<void(std::string const& writableName, std::string const& archiveName)> const&
f) = 0;
advance(std::unique_ptr<Backend>&& newWritable, RingPersist const& persist) = 0;

/**
* Marks an online-delete rotation as in progress (or completed).
*
* While in flight, a read served by the archive backend is copied
* forward into the writable backend even for ordinary
* (duplicate == false) fetches: the archive is about to be deleted,
* and a node body canonicalized into caches during the rotation
* window would otherwise survive only in RAM once the archive is
* dropped.
* Number of live generations currently in the ring.
*/
virtual std::size_t
generationCount() const = 0;

/**
* Number of live nodes copied forward out of the retiring generation during the
* current retire window (reset by beginRetire). This is the evacuation volume — the
* churn the ring pays in place of copying the whole live state every rotation — so it
* quantifies that reclamation is O(churn) rather than O(total state).
*/
virtual std::uint64_t
copyForwardCount() const = 0;

/**
* Begin/end retiring the oldest generation. While a retire is in progress, any
* read served by the retiring generation is copied forward into the writable
* backend (even ordinary reads): that generation is about to be dropped, so its
* still-live nodes must be preserved. Copy-forward is scoped to the retiring
* generation only — reads served by other sealed generations are NOT copied, which
* is what keeps evacuation O(churn) rather than O(total state).
*/
virtual void
beginRetire() = 0;
virtual void
endRetire() = 0;

/**
* Drop the oldest generation (its survivors already evacuated during the retire
* window). The generation's directory is deleted only after @p persist records the
* shortened ring, so a crash never leaves a persisted name without its backend.
*/
virtual void
setRotationInFlight(bool inFlight) = 0;
retireOldest(RingPersist const& persist) = 0;
};

} // namespace xrpl::node_store
50 changes: 33 additions & 17 deletions include/xrpl/nodestore/detail/DatabaseRotatingImp.h
Original file line number Diff line number Diff line change
Expand Up @@ -10,11 +10,13 @@
#include <xrpl/nodestore/Scheduler.h>

#include <atomic>
#include <cstddef>
#include <cstdint>
#include <functional>
#include <memory>
#include <mutex>
#include <string>
#include <vector>

namespace xrpl::node_store {

Expand All @@ -26,11 +28,11 @@ class DatabaseRotatingImp : public DatabaseRotating
DatabaseRotatingImp&
operator=(DatabaseRotatingImp const&) = delete;

// Generations are ordered oldest -> newest; the last is the writable backend.
DatabaseRotatingImp(
Scheduler& scheduler,
int readThreads,
std::shared_ptr<Backend> writableBackend,
std::shared_ptr<Backend> archiveBackend,
std::vector<std::shared_ptr<Backend>> generations,
Section const& config,
beast::Journal j);

Expand All @@ -40,10 +42,22 @@ class DatabaseRotatingImp : public DatabaseRotating
}

void
rotate(
std::unique_ptr<node_store::Backend>&& newBackend,
std::function<void(std::string const& writableName, std::string const& archiveName)> const&
f) override;
advance(std::unique_ptr<Backend>&& newWritable, RingPersist const& persist) override;

std::size_t
generationCount() const override;

std::uint64_t
copyForwardCount() const override;

void
beginRetire() override;

void
endRetire() override;

void
retireOldest(RingPersist const& persist) override;

std::string
getName() const override;
Expand All @@ -70,20 +84,22 @@ class DatabaseRotatingImp : public DatabaseRotating
void
sweep() override;

void
setRotationInFlight(bool inFlight) override;

private:
std::shared_ptr<Backend> writableBackend_;
std::shared_ptr<Backend> archiveBackend_;
// Immutable snapshot of the generation ring, ordered oldest (front) -> newest (back);
// back() is the writable backend. Replaced copy-on-write under mutex_ by advance() /
// retireOldest(); fetchNodeObject takes a shared_ptr copy and iterates it lock-free,
// so the hot read path never allocates and never blocks writers.
std::shared_ptr<std::vector<std::shared_ptr<Backend>>> ring_;
// The oldest generation while it is being retired (evacuated then dropped), else
// null. A read served by this generation is copied forward into the writable
// backend so its survivors are preserved before it is dropped; copy-forward is
// scoped to this generation only (reads from other sealed generations are not
// copied), which keeps evacuation O(churn) instead of O(total state).
std::shared_ptr<Backend> retiring_;
mutable std::mutex mutex_;

// True between SHAMapStore starting the cache-freshen phase and the
// completion of rotate(). While true, archive hits on ordinary
// (duplicate == false) fetches are copied forward into the writable
// backend; copyForwardCount_ tallies them per rotation for the
// summary line logged at swap.
std::atomic<bool> rotationInFlight_{false};
// Tally of nodes copied forward out of the retiring generation, for the summary
// line logged when the generation is dropped.
std::atomic<std::uint64_t> copyForwardCount_{0};

std::shared_ptr<NodeObject>
Expand Down
16 changes: 16 additions & 0 deletions include/xrpl/server/State.h
Original file line number Diff line number Diff line change
Expand Up @@ -5,14 +5,30 @@
#include <xrpl/rdb/SociDB.h>

#include <string>
#include <vector>

namespace xrpl {

struct SavedState
{
// Legacy two-backend fields. Retained for on-disk back-compat: a state written
// by an older (two-backend) build has only these. On read, an empty `generations`
// is reconstructed as {archiveDb, writableDb} (oldest -> newest). On write, these
// are kept in sync with the ring ends (archiveDb = generations.front(),
// writableDb = generations.back()). CAUTION: a downgraded build boots from the
// pair alone — it deletes middle-generation directories as orphans (losing any
// node whose only copy lives there) and its rotations leave DbGenerations rows
// stale; getSavedState detects that staleness (ring ends disagreeing with the
// pair) and falls back to the pair.
std::string writableDb;
std::string archiveDb;
LedgerIndex lastRotated{};

// The online_delete generation ring, ordered oldest -> newest. New nodes are
// written to generations.back() (the writable generation); reads probe
// newest -> oldest. Persisted one row per generation in the DbGenerations
// table, ordered by position, so a variable number of generations round-trips.
std::vector<std::string> generations;
};

/**
Expand Down
Loading
Loading