cobalt
An embedded LSM key/value store in Zig. Fresh writes go to a WAL-backed skiplist, flush to immutable SSTables, and space is reclaimed by leveled compaction.
An embedded LSM key/value store in Zig. Fresh writes land in a WAL-backed skiplist, then get flushed to immutable SSTables, and leveled compaction reclaims the space. It is the foundation modern databases stand on.
Overview
The modern databases you use every day almost always have the same engine shape underneath. cobalt exists so that shape stops being a black box: it implements the whole LSM write path from the log to compaction, in Zig, a language that hides not a single byte or allocation from you.
This is not another database for production but an engine taken apart into pieces where every decision is visible. Once you write the WAL, memtable, SSTable and compaction yourself, RocksDB and Cassandra stop being magic and become familiar variants of the same idea.
Why LSM, not a B-tree
Classic databases write in place, updating B-tree pages, which means random writes across the disk. LSM goes the other way: every write is an append. New data lands in an in-memory structure while older layers sit immutable on disk. That makes writes fast and sequential, because sequential writes are exactly what a disk likes most.
The cost is clear and deliberate: the same data may live in several places at once and has to be merged later. The whole art of LSM is paying that debt in the background without slowing the writes that create it.
The write path, step by step
a write hits the log first, so it survives a sudden power-off before it lands anywhere else.
fresh data sits in a sorted in-memory structure, ready for fast reads.
when memory fills up, the skiplist is flushed to an immutable, sorted on-disk file.
leveled compaction merges files across levels, drops overwritten versions and reclaims space.
The order of these steps is not accidental. The WAL comes first precisely because it guarantees durability before the data reaches volatile memory. If the order were reversed, a power failure between the memory write and the log would mean silently lost data.
Compaction is the heart that separates a working LSM from one that chokes over time. Without it, SSTable files pile up endlessly and a read has to look through more and more layers. Leveled compaction merges them across levels, throwing out versions that were overwritten and reclaiming the space taken by dead data.
The same engine shape, WAL plus memtable plus SSTable plus compaction, is what you find in RocksDB, LevelDB and Cassandra. The difference between them is mostly tuning of the same building blocks, not a different architecture.
The end of the black box
The biggest payoff of cobalt is not the store itself but that big databases stop being a mystery. Once you have walked the whole path from appending to the log to reclaiming space in compaction, you know exactly why LSM is fast on writes and where it pays for it on reads.
The engine layers and where they also live
| Layer | Role | Also found in |
|---|---|---|
| WAL | write durability | every serious database |
| Memtable (skiplist) | fresh data in memory | RocksDB, LevelDB |
| SSTable | immutable on-disk files | Cassandra, LevelDB |
| Compaction | space reclaim, merging | RocksDB, Cassandra |
More projects
More work from the same category - see how we tackle similar challenges.
Have a similar project?
Get in touch - a quote is free and comes back within an hour.



