My apologies for the long time no see, but there was a tectonic shift (colloquially known as a pivot) in our company's positioning. While our website is still all about gamedev (we're working on a redesign, but these things take time), I am happy to announce that we've shifted from developing our own game engine to plain B2B, with an emphasis on (a) DBMS and (b) CI infrastructure for developing bulletproof-by-design products.

It took quite a bit of time to get here, but we've managed to design a few interesting things along the way. So now, before continuing to publish chapters of "Efficient C++ Programming for Modern 64-bit CPUs", I'm going to make a detour into publishing - and inviting your comments on - a radically unusual architecture for an upcoming RDBMS engine.


MECHLOVE Blueprint - 1/7. General: Pretty Much Classical RDBMS at Heart, with Each and Every Component Rewritten


MECHLOVE Blueprint - 1/7. General: Pretty Much Classical RDBMS at Heart, with Each and Every Component Rewritten - Image 1

Preamble

  1. At this point, what we’re doing is publishing a STATEMENT OF INTENT; in other words, we do NOT have MECHLOVE yet (though we’re already working on implementation), but we want to attract constructive criticism from the community to validate our (often rather unusual) approaches. As a result, any performance claims are by necessity merely performance targets/guesstimates. 
  2. This document isn’t a full blueprint; further chapters are ready, and we plan to publish one chapter per week.

Upcoming Parts:

  • Radical MVCC
  • Replay-Based Rebasing OCC (Re2OCC)
  • MECHLOVE Blueprint - 2/7. Storage (ZFS-style) and WAL (template-based logical with physical hints)
  • MECHLOVE Blueprint - 3/7. Bufferpools: In-Memory CoW, Latch-less Cache-friendly SIMD-friendly Layouts with Built-In Versioning
  • MECHLOVE Blueprint - 4/7. Multiple Writers: Drastic Reduction of Wasted Work and Starvation-Free Guarantees by Replay-Based Rebasing OCC
  • MECHLOVE Blueprint - 5/7. SQL Compiler, Execution Engine, and Schema Versioning
  • MECHLOVE Blueprint - 6/7. Replication and HA
  • MECHLOVE Blueprint - 7/7. Supported Platforms, Editions, and Roadmap 

MECHLOVE

MECHLOVE stands for “MECHanical LOVE”, and we consider MECHLOVE to be the highest form of mechanical sympathy. The idea is to embrace the properties of modern hardware, instead of trying to fight them. As modern CPUs actively hate pointer chasing and MESI traffic, we’ll avoid them as much as possible; as SSDs are still in love with linear scanning, we’ll embrace it too, and so on and so forth. 

What MECHLOVE is About

Classical RDBMS, implemented accounting for realities of 2026+

We’re taking classical RDBMS architecture (with WAL, bufferpools, pages, MVCC, etc.) - and (being in LOVE with CPU mechanics) are implementing this classical RDBMS from scratch in a hardware-friendly manner. Improvements include being lockless (even atomic RMWs on the mainstream read path are avoided), cache-friendly, SIMD-oriented, and so on. At the same time, classical RDBMS architecture still ensures classical RDBMS properties such as full-scale ACID, bounded crash recovery times, PITR recovery, a solid foundation for replication and HA, and so on and so forth, from the very beginning. To clarify what “Classical RDBMS” is in our parlance:

  • two durability domains: Data and WAL;
  • WAL is written+fsync()-ed eagerly; Data is written lazily;
  • it makes checkpoints necessary (and for operational reasons, in modern RDBMS checkpoints have to be fuzzy, without stopping the world);
  • crash recovery is done via re-applying a portion of WAL.

In addition, while not strictly required by the model above, we decided to keep pages.

Optimized for Plenty of RAM

While we do not require that the whole DB resides in RAM, we optimize for having plenty of RAM, sufficient for caching all the active working set in RAM. In addition, we are perfectly ready to enforce that some meta-information (such as ZFS-like uberblocks, more on it in Part 2) is always in RAM.

Optimized for Modern CPUs

If we look at the landscape of popular RDBMS (from MySQL/Postgre/SQLite to Oracle/DB2/MS SQL), we’ll see that all of them were designed before 2000; since then, CPUs have changed drastically, while dominant RDBMS architectures have stayed the same. Among other things, (a) pointer-chasing, which wasn’t an issue 30 years ago, became a dominant cost in lock-less programming (and in spades for RDBMS), and (b) due to modern CPUs being heavily OoO (with typical RIPC reaching 4 on non-RDBMS type of load) and astonishingly smart modern compilers, even an indirection via a function pointer becomes highly visible performance-wise.

Optimized for Integrity and Serializable Snapshot Isolation

As our Chief Architect also used to be an architect of stock exchange software for several countries, and of another system that processed billions of write transactions and hundreds of millions of euros per year, her stance on integrity is extremely hard. In particular, we do NOT condone error-prone programming styles that have to rely on techniques such as SELECT FOR UPDATE to avoid races, and on lock ordering to avoid deadlocks. Our position is that RDBMS shall be that fast even at the highest-possible Serializable Snapshot Isolation (SSI) level, so that developers will never need to risk integrity for performance (and will never ever need to resort to lower isolation levels due to performance considerations). In fact, we plan to support SSI as the only isolation level (and make it completely transparent to the application level, too). 

OLTP first

While overall MECHLOVE architecture is flexible enough to support pretty much any load, at first, we’ll be paying special attention to OLTP loads. In particular, at this stage we’re leaving aside OLAP-oriented optimizations such as columnar storage and OLAP-style vectorized execution (note that while we do use SIMD, we’re using it in a substantially different manner). 

Open Source

Most of MECHLOVE will be open-sourced (as Community Edition); only enterprise-oriented features such as HA or encrypt-at-rest will stay behind the paywall of Enterprise Edition, and all the performance features discussed in these blueprints will be open-sourced. We plan to open-source MECHLOVE core as soon as we have something coherent (we don't want to publish it right now to avoid being criticized for essentially WIP-grade code; benchmarking such code will be especially ugly for the project).

What MECHLOVE is NOT

Now, with all the modern trends trying to disrupt the RDBMS landscape, we have to clearly state what MECHLOVE is NOT. There are quite a few modern trends which we consider way too radical and/or inapplicable for our purposes. 

NOT a NoSQL DB

Unlike lots of systems out there, MECHLOVE is a classical SQL-oriented RDBMS with strict multi-row ACID guarantees. Actually, it is multi-row ACID, which is traditionally an Achilles' heel of NoSQL, and trying to build any kind of transactional system without it is controversial to say the least. 

NOT a "Cloud-Native" Heavily Distributed System (Raft/Paxos/Distributed SQL/etc.)

These days, everybody and their dog is rooting for heavily distributed systems relying on consensus protocols such as Raft and Paxos. While these protocols are theoretically beautiful, we do not feel that they are optimal for real-world deployments. In particular, their whole approach tries to solve three quite different problems together:

  • High Availability (HA): we feel HA should be handled by much simpler means (specifically, via replication + delayed replies + deterministic logic to handle cross-datacenter 2- and 3-node configurations, see Part 6 for further discussion);
  • Reader scaling: this is trivially handled via CQRS (traditionally implemented via replication);
  • Writer scaling. If you go beyond embarrassingly shardable loads, distributed systems do NOT really handle write scaling well; moreover, this isn't really a flaw of consensus protocols, but a fundamental restriction. Without sharding, active-active writers have never scaled horizontally, and are very unlikely to scale, ever. Worse than that - for consensus protocols, effectively having one log at each point makes it a chokepoint by definition. This, in turn, makes the scaling of a single node of paramount importance. 

Trying to kill several birds with one shot isn't inherently bad, but let’s not forget these distributed systems come with a huge cost; most importantly, they tend to be MUCH more resource-consuming than classical RDBMS. Let’s just compare Hekaton reaching tpmC of 1.25M on a single 80-core server back in 2014 with CockroachDB reaching tpmC of 1.68M on whopping 81 boxes, each node having 36 vCPUs, and Yugabyte DB reaching tpmC of 1M on 75 nodes, each node having 48 vCPUs; that’s at least a 10x penalty for heavily-distributed systems (that's even after we account for 2-datacenter HA). While this is certainly nice for hyperscalers charging per vCPU, it is not so nice for businesses using them. Add on top of it the operational complexity of dealing with 20 nodes instead of just 2-3, and you’ll get the idea why our Chief Architect (with her real-world experience described above) is not too fond of heavily distributed "cloud-native" systems. 

NOT an LSM-based system

LSM-based systems are once again very nice in theory, but are pretty poorly suited to typical OLTP transactional loads such as TPC-C. Most of them are not able to handle regular SQL, and those that can (such as MySQL+RocksDB) have no published tpmC numbers. Exceptions are the LSM-based distributed systems we already discussed (Cockroach and Yugabyte), but as we’ve seen above, they suffer from huge performance penalties.

As a side note, one potential way to see MECHLOVE is as an LSM-like architecture with an in-memory compaction (see "materialization" in the "Radical MVCC" document); we're not going to go into a theoretical debate about whether it is a valid taxonomy; we just want MECHLOVE to scale and perform well.

NOT Shared-Nothing with Manual Sharding

RDBMS such as VoltDB enforce a shared-nothing model, which is (again!) very nice in theory; however, once again, it comes with a huge drawback: VoltDB uses stop-the-world for inter-shard transactions, so while it can reach pretty high numbers for embarrassingly shardable TPC-C-like loads, it requires manual sharding and is unlikely to scale well into less-embarrassingly-trivially-sharded OLTP workloads.

NOT a Fully In-Memory DB

Quite a few modern databases, such as Hekaton (now named “Microsoft SQL Server In-Memory OLTP”), VoltDB, Silo, and HyPer, are in-memory only. This means higher RAM requirements and a lack of graceful degradation when the “100% DB fit in memory” requirement is no longer satisfied. In addition, we feel that using RAM for storing almost-never-read historical tables (which constitute the majority of most real-world OLTP systems out there, see e.g. TPC-C) is a bad idea, and that this very RAM can be utilized in a much more efficient manner (see e.g. our “hypercache” in Part 3).

Some of the experimental databases go even further, abolishing pages entirely, and relying on lower-granularity in-memory data structures. This, however, way too often leads to pointer chasing, which is devastating to the performance of modern CPUs. In contrast, we don't think pages should be abolished; instead, we treat pages as our primary unit of allocation and versioning. Moreover, it plays nicely with the all-important concept of pre-allocating everything. And with this in mind, we don’t think we’ll get enough performance benefits from keeping our pages 100% in memory (hey, you still need checkpoints for in-memory data).

NOT an Experimental/Research DB

We are NOT trying to make yet another experimental/research DB like Silo or HyPer. From the outset, we’re aiming to build a production RDBMS with all the standard features, such as full ACID compliance, crash recovery, O(1) ADD COLUMN, and heterogeneous replication, accounted for (at the very least, we know how we’ll implement each and every one of them without patching the core). That being said, we’re aiming to match or exceed Silo and HyPer performance-wise.

MECHLOVE: General Architecture

At a very high level, we’re using a classical architecture with WAL (Write-Ahead Log) written eagerly (and fsync()-ed before confirming the transaction as committed to the user), and data stored in pages which are written asynchronously (and with checkpoints being fuzzy). This provides potential for high performance alongside strict ACID guarantees.

Back to Codd, Kung, and Robinson

The totality of data in a data bank may be viewed as a collection of time-varying relations.

– E.F. Codd

Our overall philosophy is that we’re working with data sets, and rather simple math involving these data sets. Just as one example, instead of modifying pages on the fly, our writing thread keeps local changes as a thread-local write set and interposes it over the immutable MVCC snapshot whenever it tries to read. As another example, instead of issuing global SILocks, we’re directly comparing the candidate read-set (expressed in terms of predicates to account for phantom reads) and write-set with the write-set of the already-committed transaction.

We feel that this approach is actually closer to the original works of Codd and Kung&Robinson than most existing RDBMS implementations; after all, Codd didn’t write in terms of locks, and Kung&Robinson’s OCC mentioned them only as a last-resort measure to avoid starvation.

Our Restrictions

Our MECHLOVE engine comes with two major restrictions: 

  • We strictly require each table to have a Primary Key (PK). This is one de facto firm requirement for any sane relational database design anyway (BTW, original Codd’s relational model strictly requires at least one “candidate key”, which means it should always be possible to make a PK out of it).
  • We require that all the writing transactions are “One-Shot Transactions”. In other words, we require that all the writing DB transactions are well-known in advance. We will support updating transactions on the fly (at first, simply using LLVM to compile them to .so/.DLL under the hood), but ad-hoc (user-driven) writing transactions will be out of reach at least for Community Edition. Enterprise Edition will support ad-hoc writing transactions, but they will still carry a huge warning sign: long-running ad-hoc transactions will hurt performance greatly (as they do in any RDBMS anyway). To put it into perspective, our “One-Shot Transactions” restrictions are conceptually similar to those of stored procedures (on the stricter side, with an outright prohibition on the network/IO). 

Our Benefits

In exchange for these restrictions, developers will get:

  • No worries about concurrency whatsoever. As noted above, the only isolation level we’ll support is Serializable Snapshot Isolation. Look, ma, neither data races nor deadlocks! It is conceptually similar to Postgres’ SSI, but with One-Shot Transactions, we can implement Re2OCC, which “rebases” transactions instead of retrying them, which in turn (a) avoids throwing out lots of work, (b) doesn’t require app-level handling, and (c) in a Re2OCC-SF variation, formally guarantees against indefinite starvation. 
  • Extreme SQL performance. This is actually much more elaborate than it is usually implied when speaking about performance benefits of One-Shot Transactions. First of all, one-shot transactions enable Re2OCC (more on it in a separate document), which in turn greatly reduces the waste of OCC retries in all-important use cases. Second, it allows analysis of the whole transaction (in contrast to per-statement analysis), which we will use to improve performance further. And third, a lot of runtime indirections can be eliminated if compiling them in. Usually, only the 3rd factor out of those listed above is considered when speaking about the benefits of One-Shot Transaction; as we see above, it is much wider than that, plus, as it directly follows from Amdahl’s Law, even the third factor will produce much better performance improvements percentage-wise as soon as we reduce other costs (such as lock costs and pointer-chasing costs), as described below.
  • Extreme performance even when using ORM. One-shot transactions are probably the only way to compile ORM into sensible SQL (and then to “ideal” native code as described above), avoiding well-known ORM artifacts such as N+1 queries and over-fetching. While it is still in the research phase (and technically belongs to our other DB-related project - NYUORM), if we manage to pull it off, it will be huge; the ability to produce highly efficient RDBMS code without thinking in terms of SQL would be a major improvement in developer productivity.

The Devil is in the Details: Improvements Compared to Classical RDBMS

While we do follow classical RDBMS architecture (data storage + WAL, and even embrace pages and bufferpools) - we redesign each and every aspect of them. We’ll discuss improvements over classical RDBMS in detail in subsequent Parts, but here is a very high-level overview of our improvements:

  • ZFS-like integrity features including disk-level CoW, end-to-end checksumming, protection against bit rot, and [Enterprise Edition] self-healing and Data-At-Rest-Encryption. As a side benefit, disk-level CoW greatly simplifies crash recovery (we need neither undo nor DWB/FPW). Note that unlike ZFS, we use versioned disk-level CoW to support MVCC seamlessly;
  • Radical MVCC. Unlike the vast majority of RDBMS in the wild, we do not write transaction updates into pages until the transaction is known to be valid and has obtained its own CSN. This gives us an extremely clean versioning model, and lazy consolidatable updates of in-memory pages too;
  • In-memory versioned CoW (with versioned SIMD-friendly patches; see Part 3 for details). This is probably our most important “secret weapon” to achieve the highest possible performance. Very briefly, it provides several all-important benefits. Not only do we (a) avoid page locks/latches on the main execution paths and (b) avoid atomic RMWs on the mainstream read path (which is similar to Bw-tree and Hekaton), we also manage to (c) avoid pointer chasing and, more generally, have CPU-cache- and SIMD-friendly layouts. For further discussion, see Part 3;
  • Abolishing ARIES-style “physiological WAL” and replacing it with logical WAL (with certain improvements; see Part 2 for details). This, in turn, opens the door to redesign page layouts to be latch-free, cache-friendly, and SIMD-friendly. NB: to avoid any doubt, our logical WAL is idempotent, with WAL frames semantically being along the lines of “UPDATE T SET F=? WHERE PK=?”;
  • Re2OCC. As big fans of lock-free processing, we’re naturally big fans of OCC too; however, classical OCC wastes a lot of work under contention. Our Re2OCC “rebases” conflicting transactions; moreover, it uses conflicting writer sets to augment already-obtained results whenever possible, which speeds up a lot of real-world use cases. We use Re2OCC-SF, a variation with strict guarantees against starvation (and without too much performance loss).

Key Invariant

If trying to summarize the key invariant of MECHLOVE in one sentence, it would be something along the lines of:

Primary database state is the ledger of the CSNs with corresponding write-sets; 

everything else (including WAL, pages, conflict detection, and rebasing) is derived from this ledger.

Note that this doesn’t mean we’re a classical LSM system; we materialize this state into classical pages within milliseconds, so searches don’t need to scan the whole history, and we truncate the ledger as soon as it's no longer needed for any of the purposes listed above. OTOH, a discussion on "whether MECHLOVE can be seen as an LSM system with an almost-immediate in-memory compaction" is beyond our current scope.

Let’s further note that this primary state is obviously not durable until it is written to WAL and fsync()'ed; however, ACID Durability is a client-side property and can be (and is) ensured by simply delaying the reply to the client until the respective transaction is fsync()'ed. In a similar manner, violations of linearizability are prevented by delaying reply to the client until the pages of the respective transaction are materialized. 

Performance Target

According to our very preliminary mini-research (extrapolating existing data to some standardized hardware), for full-ACID-compliant TPC-C TPS on a 96-core Zen5 box, an order-of-magnitude guesstimate looks as follows: 

RDBMS

Postgre / MySQL

Oracle / DB2 / MSSQL

Hekaton

MECHLOVE

TPC-C TPS

30-50k

50-100k

300k

1M

While this may look overly ambitious, let’s note that 1M TPS on a 96-core box means that we have mere 10K TPS per core, or a whopping 300’000 CPU cycles per transaction. Moreover, our theoretical-limit estimates show that even 30’000 CPU cycles per transaction are not proven unreachable, so at 1M TPS we still have about a 10x reserve relative to the theoretical limit.

Why Is It Possible?

RDBMS mentioned above were designed before 2000 (the only exception being Hekaton, designed circa 2012), and we’re now in 2026. And now we have an inherent advantage of having knowledge of (a) the environment where the RDBMS will be running (see above re. modern CPUs and plenty of RAM), and (b) new solutions which were simply unknown back in 199x. Examples of such drastic changes in technology, which we’ll be using/relying on, include:

  • cache-friendly layouts (CSB+ trees were invented in 2000; [disclaimer: we won’t be using CSB+ trees as such, but we will use cache-friendly layouts which are also SIMD-friendly and latch-free]);
  • seqlock (invented in 2002);
  • 64-bit CPUs became mainstream (starting with AMD Opteron in 2003); while modern RDBMS were ported to support 64-bit space pretty soon, optimizing for Plenty of RAM is still in progress;
  • QSBR (invented in 2002, got its name in 2006);
  • ZFS (released in 2005), implementing, in particular, on-disk CoW;
  • philosophy of in-memory CoW (popularized between 2007-2015);
  • related but distinct latch-free philosophy (for example, Bw-tree was invented circa 2010 and published in 2013 [disclaimer: we won’t be using Bw-trees as such, but we will achieve latch-free and atomic-RMW-free on mainstream paths, while also being cache- and SIMD-friendly]);
  • C++ templates matured sufficiently to become usable in app-level code (circa C++11);
  • SIMD-friendly layouts (for example, Advanced Radix Tree was invented in 2013 [disclaimer: we won’t be using ART as such, but we will use SIMD-friendly layouts which are also latch-free and cache-friendly]);
  • cache- and SIMD-friendly Swiss Tables and F14 were invented in 2017 and 2019, respectively;
  • C++ constexpr matured sufficiently to become usable in app-level code (C++14 - C++20);
  • Intel’s and AMD’s clarification that, given proper alignment, ZMM loads are immune to intra-CAS torn reads - issued in 2022;
  • versioned-pointers (verlib was released in 2024 [disclaimer: we invented our own versioned-pointer independently, and unlike verlib’s one, it is cache-friendly and SIMD-friendly too]);
  • Radical MVCC and Re2OCC, our own inventions (or adaptations of existing approaches to our purposes) of 2026. 

Using some of the techniques above, Hekaton was able to achieve a 3x performance improvement over the best classical RDBMS. However, while their data structures are latch-free, they are neither cache-friendly nor SIMD-friendly. It seems that we should be able to have pages which are (a) latch-free (and even atomic-RMW-free for reads), (b) cache-friendly, and (c) SIMD-friendly, all at the same time; see Part 3 for details. This, in turn (combined with hyper-cache described in the same Part 3), should provide a very significant performance boost. In particular, most of the time spent by engines such as Hekaton is pointer-chasing (in an OLTP context, most memory jumps take about 200 CPU cycles each), and being cache- and SIMD-friendly is expected to reduce such jumps by factors of 3x to 10x. 

One Hell of an Async Design

It currently seems that (using a combination of Radical MVCC and OCC) we managed to:

  • avoid shared RMWs on the mainstream read paths;
  • have only one fundamental CAS in the system (transactions trying to obtain CSN and publish their respective write-set, both by the same CAS). Having the only CAS in the whole RDBMS mainstream operation is extremely difficult to beat (especially as this ordering is crucial for the RDBMS logic);
  • keep the vast majority of the code sync-free, so it can be reasoned about in terms of good ol' state machines and set math.

There is, of course, no magic involved; concurrent transactions still have to interact (at least to validate); what we’re essentially playing with is increasing the granularity of the synchronizations. And (both performance-wise and complexity-wise) having a few coarse-grained synchronization points (such as “we finished a transaction, here is its now-immutable write-set”) is orders of magnitude better than having zillions of fine-grained ones (such as “we modified a row, make sure to read an updated one”).

To Be Continued...

As noted above, this is certainly NOT the whole blueprint.

The following documents are ready, and will be published in the due course:

  • Radical MVCC
  • Replay-Based Rebasing OCC (Re2OCC)
  • MECHLOVE Blueprint - 2/7. Storage (ZFS-style) and WAL (template-based logical with physical hints)
  • MECHLOVE Blueprint - 3/7. Bufferpools: In-Memory CoW, Latch-less Cache-friendly SIMD-friendly Layouts with Built-In Versioning
  • MECHLOVE Blueprint - 4/7. Multiple Writers: Drastic Reduction of Wasted Work and Starvation-Free Guarantees by Replay-Based Rebasing OCC
  • MECHLOVE Blueprint - 5/7. SQL Compiler, Execution Engine, and Schema Versioning
  • MECHLOVE Blueprint - 6/7. Replication and HA
  • MECHLOVE Blueprint - 7/7. Supported Platforms, Editions, and Roadmap