DuckDB Ducklake
github.com/duckdb
[6 comments hidden]
[3 comments hidden]
[2 comments hidden]
[hidden]
[2 comments hidden]
https://duckdb.org/2025/05/19/the-lost-decade-of-small-data....
[3 comments hidden]
[3 comments hidden]
For example I think a lakehouse is a bad choice for standard enterprise BI type analytics - you've got no column or row access controls, and no column masking. I don't see how this could ever be bolted on to the bucket and catalog.
https://www.tomwphillips.co.uk/2026/08/the-benefits-of-data-...
[2 comments hidden]
[4 comments hidden]
[3 comments hidden]
Much faster, but adds a dependency... that you would have added anyway with database based catalogs (that are not the only kind of catalogs)
[8 comments hidden]
[6 comments hidden]
[5 comments hidden]
[3 comments hidden]
[2 comments hidden]
[hidden]
As someone who's personal spare time tinkering project is writing DBMS + storage engine in Rust I will say the language choice or it's memory safety is not that consequential.
If you're doing anything dealing with an actual storage engine you will be writing a bunch of unsafe code and using raw pointers all over the place. My hypothesis is once the guts are in place then I can go a lot faster with Rust but it's very much swimming against the tide and it is not magically safe out of the box once you are dealing with raw storage and paging.
But nonetheless if you go around the industry and do a count of commercial and Open source DBMS systems that have any kind of significant adoption it'll be about 95% C/C++ even if you filter to recent systems and there are legitimate reasons for that being a common choice.
TL:DR The hard part of a DBMS implementation is not picking the language
[hidden]
However, DuckLake doesn't need to be implemented in C++ at all, I believe that whatever fits the engine that will be performing the reads/writes in Parquet and the connectors for the DBMSs should be a good fit.
For example https://github.com/borchero/ducklake-sdk - Rust https://github.com/datafusion-contrib/datafusion-ducklake - Rust https://github.com/motherduckdb/ducklake-spark - Scala https://github.com/brikk/trino-ducklake - Kotlin
Disclaimer: I'm the lead developer of DuckLake.
smithclay[7 comments hidden]
There's a cool alternate rust/datafusion ecosystem initiative going on at https://github.com/datafusion-contrib/datafusion-ducklake, and think the Quack protocol opens up a lot of cool possibilities too.
If you need an idea for what do do with ducklake: recommend throwing all of your agent traces in it.
eddietejeda[3 comments hidden]
For context, we previously built custom catalogs optimized for specific use cases. But they were hard to maintain, especially as requirements changed, and Apache Iceberg was too heavy for our specific low-latency work.
Since Ducklake is only a spec, we implemented datafusion-ducklake, and it performs as well as any custom or specialized catalog we built. We use Postgres as the catalog store, and it does not get much simpler than that: a transactional database for transactional data.
Plus, it gives us a clear spec for implementing complex parts like time travel, snapshots, etc.
It's been a godsend.
We welcome and encourage contributors!
prpl[2 comments hidden]
eddietejeda[hidden]
Here is our write up on the Ducklake blog: https://ducklake.select/2026/07/29/bringing-ducklake-to-data...
wodenokoto[3 comments hidden]
I always thought the catalogue was a duckdb file. E.g, data lives in partitioned parquet files, but which parquet files are current or soft deleted, etc, etc, is managed in a duckdb data file.
However, looking at https://ducklake.select/, it seems the catalogue lives in PostgresSQL - so it is not really a ducklake, but a postgresslake.
The more you know.
paragraft[hidden]
pdet[hidden]
There are ofc some quirks from the SQL supported on each DBMS. Hence, the DuckLake extension from DuckDB currently supports DuckDB, SQLite, Postgres, DuckDB + Quack and MySQL. With MySQL being in a rather experimental state.
Disclaimer: I'm the lead developer in DuckLake.