Issue summary
Any WAL-enabled write on a DB opened with (unordered_write = true, manual_wal_flush = true, two_write_queues = false) deadlocks. The thread that issues the write blocks forever in pthread_mutex_lock inside DBImpl::WriteToWAL, waiting on wal_write_mutex_, which the same thread already holds because DBImpl::ConcurrentWriteGroupToWAL acquired it before calling WriteToWAL.
Simple test
Adding this test to rocksdb/db/db_write_test.cc can see it failing because of deadlock:
TEST_F(DBWriteTest, UnorderedWriteManualWalFlushDoesNotDeadlock) {
Options options = CurrentOptions();
options.unordered_write = true;
options.manual_wal_flush = true; // two_write_queues stays false (default)
DestroyAndReopen(options);
std::promise<Status> p;
auto f = p.get_future();
std::thread writer([&] { p.set_value(db_->Put(WriteOptions(), "k", "v")); });
if (f.wait_for(std::chrono::seconds(10)) == std::future_status::timeout) {
writer.detach(); // stuck in pthread_mutex_lock; can't join
FAIL() << "Put deadlocked (unordered_write + manual_wal_flush + !two_write_queues)";
}
writer.join();
ASSERT_OK(f.get());
}
Mini-repro of the deadlock with DB put
#include <chrono>
#include <cstdio>
#include <memory>
#include <string>
#include <thread>
#include "rocksdb/db.h"
#include "rocksdb/options.h"
int main(int argc, char** argv) {
const std::string db_path =
argc > 1 ? argv[1] : "/tmp/rocksdb-manual-wal-flush-deadlock";
rocksdb::Options options;
options.create_if_missing = true;
options.unordered_write = true;
options.manual_wal_flush = true;
// two_write_queues is false by default; that's the third condition.
std::unique_ptr<rocksdb::DB> db;
auto s = rocksdb::DB::Open(options, db_path, &db);
if (!s.ok()) {
std::fprintf(stderr, "DB::Open failed: %s\n", s.ToString().c_str());
return 1;
}
std::fprintf(stderr, "DB opened. About to Put (this should hang)...\n");
std::thread watchdog([]() {
std::this_thread::sleep_for(std::chrono::seconds(10));
std::fprintf(stderr,
"*** Put has been blocked for 10 seconds -- this is the "
"deadlock. Kill with SIGABRT to get a stack trace.\n");
std::abort();
});
watchdog.detach();
s = db->Put(rocksdb::WriteOptions(), "key", "value");
std::fprintf(stderr, "Put returned: %s (unexpected, no deadlock detected)\n",
s.ToString().c_str());
return 0;
}
References
Issue summary
Any WAL-enabled write on a
DBopened with(unordered_write = true, manual_wal_flush = true, two_write_queues = false)deadlocks. The thread that issues the write blocks forever inpthread_mutex_lockinsideDBImpl::WriteToWAL, waiting onwal_write_mutex_, which the same thread already holds becauseDBImpl::ConcurrentWriteGroupToWALacquired it before callingWriteToWAL.Simple test
Adding this test to
rocksdb/db/db_write_test.cccan see it failing because of deadlock:Mini-repro of the deadlock with DB put
References
db/db_impl/db_impl_write.cc:1817-1866(ConcurrentWriteGroupToWAL, acquires the mutex): https://github.com/facebook/rocksdb/blob/v10.9.1/db/db_impl/db_impl_write.cc#L1817db/db_impl/db_impl_write.cc:1653-1698(WriteToWAL, re-acquires the same mutex): https://github.com/facebook/rocksdb/blob/v10.9.1/db/db_impl/db_impl_write.cc#L1653db/db_impl/db_impl_write.cc:1672(the guard that misses the third caller): https://github.com/facebook/rocksdb/blob/v10.9.1/db/db_impl/db_impl_write.cc#L1672db/db_impl/db_impl_write.cc:513-538(WriteImpldispatch toWriteImplWALOnlyunderunordered_write): https://github.com/facebook/rocksdb/blob/v10.9.1/db/db_impl/db_impl_write.cc#L513