- Aerospike
- Akamas
- AlloyDB
- Antithesis
- APOLLO
- Aurora DSQL
- Berkeley DB
- BlazingDB
- Brytlyt
- Chaos Mesh
- Chronon
- ClickHouse
- Confluent
- CouchDB
- CrocodileDB
- Databricks
- Datometry
- DB9
- Debezium
- Dolt
- Dremio
- DSQL
- DVMS
- EraDB
- eXtremeDB
- Fauna
- Featureform
- Firebolt
- Fluree
- FoundationDB
- Gel
- Google Spanner
- Greenplum
- HarperDB
- HorizonDB
- Iceberg
- InfluxDB
- kdb
- ksqlDB
- LanceDB
- Litestream
- Malloy
- MariaDB
- MemSQL
- Milvus
- MonetDB
- Mooncake
- Multigres
- Napa
- NoisePage
- NuoDB
- OpenDAL
- OtterTune
- OxQL
- Pinecone
- Pixeltable
- Polaris
- PostgreSQL
- Qdrant
- QuasarDB
- RavenDB
- RelationalAI
- RocksDB
- RonDB
- SalesForce
- ScyllaDB
- SingleStore
- sled
- Smooth
- SpacetimeDB
- SpiceDB
- SplinterDB
- SQL Server
- SQLite
- Stardog
- Striim
- Swarm64
- Technical University of Munich
- TerminusDB
- TigerBeetle
- TimescaleDB
- TonicDB
- Trino
- Turso
- Velox
- VillageSQL
- VoltDB
- Weaviate
- XTDB
- YugabyteDB
- AirFlow
- Alibaba
- Anna
- ApertureDB
- Arrow
- Azure Cosmos DB
- BigQuery
- Bodo
- Cassandra
- Chroma
- Citus
- CockroachDB
- Convex
- CrateDB
- Daft
- DataFusion
- Datomic
- dbt
- Delta Lake
- Doris
- Druid
- DuckDB
- EdgeDB
- Exon
- FASTER
- FeatureBase
- Feldera
- Floe
- Fluss
- Gaia
- GlareDB
- GoogleSQL
- GreptimeDB
- Heron
- Hudi
- Impala
- Jepsen
- Kinetica
- Lakebase
- LeanStore
- LMDB
- MapD
- Materialize
- Microsoft SQL Server
- Modin
- MongoDB
- MotherDuck
- MySQL
- Neon
- Noria
- OceanBase
- Oracle
- Oxla
- ParadeDB
- Pinot
- PlanetScale
- PostgresML
- PRQL
- QMDB
- QuestDB
- Redshift
- RisingWave
- Rockset
- rqlite
- Samza
- Sentry
- Sirius
- SLOG
- Snowflake
- Spice.ai
- Splice Machine
- SQL Anywhere
- SQLancer
- SQream
- StarRocks
- Summingbird
- Synnada
- TeraData
- TiDB
- TileDB
- Tokutek
- TopK
- turbopuffer
- Umbra
- Vertica
- Vitesse
- Vortex
- WiredTiger
- Yellowbrick
- Aerospike
- Alibaba
- Antithesis
- Arrow
- Berkeley DB
- Bodo
- Chaos Mesh
- Citus
- Confluent
- CrateDB
- Databricks
- Datomic
- Debezium
- Doris
- DSQL
- EdgeDB
- eXtremeDB
- FeatureBase
- Firebolt
- Fluss
- Gel
- GoogleSQL
- HarperDB
- Hudi
- InfluxDB
- Kinetica
- LanceDB
- LMDB
- MariaDB
- Microsoft SQL Server
- MonetDB
- MotherDuck
- Napa
- Noria
- OpenDAL
- Oxla
- Pinecone
- PlanetScale
- PostgreSQL
- QMDB
- RavenDB
- RisingWave
- RonDB
- Samza
- SingleStore
- SLOG
- SpacetimeDB
- Splice Machine
- SQL Server
- SQream
- Striim
- Synnada
- TerminusDB
- TileDB
- TonicDB
- turbopuffer
- Velox
- Vitesse
- Weaviate
- Yellowbrick
- AirFlow
- AlloyDB
- ApertureDB
- Aurora DSQL
- BigQuery
- Brytlyt
- Chroma
- ClickHouse
- Convex
- CrocodileDB
- DataFusion
- DB9
- Delta Lake
- Dremio
- DuckDB
- EraDB
- FASTER
- Featureform
- Floe
- FoundationDB
- GlareDB
- Greenplum
- Heron
- Iceberg
- Jepsen
- ksqlDB
- LeanStore
- Malloy
- Materialize
- Milvus
- MongoDB
- Multigres
- Neon
- NuoDB
- Oracle
- OxQL
- Pinot
- Polaris
- PRQL
- QuasarDB
- Redshift
- RocksDB
- rqlite
- ScyllaDB
- Sirius
- Smooth
- Spice.ai
- SplinterDB
- SQLancer
- Stardog
- Summingbird
- Technical University of Munich
- TiDB
- TimescaleDB
- TopK
- Turso
- Vertica
- VoltDB
- WiredTiger
- YugabyteDB
- Akamas
- Anna
- APOLLO
- Azure Cosmos DB
- BlazingDB
- Cassandra
- Chronon
- CockroachDB
- CouchDB
- Daft
- Datometry
- dbt
- Dolt
- Druid
- DVMS
- Exon
- Fauna
- Feldera
- Fluree
- Gaia
- Google Spanner
- GreptimeDB
- HorizonDB
- Impala
- kdb
- Lakebase
- Litestream
- MapD
- MemSQL
- Modin
- Mooncake
- MySQL
- NoisePage
- OceanBase
- OtterTune
- ParadeDB
- Pixeltable
- PostgresML
- Qdrant
- QuestDB
- RelationalAI
- Rockset
- SalesForce
- Sentry
- sled
- Snowflake
- SpiceDB
- SQL Anywhere
- SQLite
- StarRocks
- Swarm64
- TeraData
- TigerBeetle
- Tokutek
- Trino
- Umbra
- VillageSQL
- Vortex
- XTDB
- Aerospike
- AlloyDB
- APOLLO
- Berkeley DB
- Brytlyt
- Chronon
- Confluent
- CrocodileDB
- Datometry
- Debezium
- Dremio
- DVMS
- eXtremeDB
- Featureform
- Fluree
- Gel
- Greenplum
- HorizonDB
- InfluxDB
- ksqlDB
- Litestream
- MariaDB
- Milvus
- Mooncake
- Napa
- NuoDB
- OtterTune
- Pinecone
- Polaris
- Qdrant
- RavenDB
- RocksDB
- SalesForce
- SingleStore
- Smooth
- SpiceDB
- SQL Server
- Stardog
- Swarm64
- TerminusDB
- TimescaleDB
- Trino
- Velox
- VoltDB
- XTDB
- AirFlow
- Anna
- Arrow
- BigQuery
- Cassandra
- Citus
- Convex
- Daft
- Datomic
- Delta Lake
- Druid
- EdgeDB
- FASTER
- Feldera
- Fluss
- GlareDB
- GreptimeDB
- Hudi
- Jepsen
- Lakebase
- LMDB
- Materialize
- Modin
- MotherDuck
- Neon
- OceanBase
- Oxla
- Pinot
- PostgresML
- QMDB
- Redshift
- Rockset
- Samza
- Sirius
- Snowflake
- Splice Machine
- SQLancer
- StarRocks
- Synnada
- TiDB
- Tokutek
- turbopuffer
- Vertica
- Vortex
- Yellowbrick
- Akamas
- Antithesis
- Aurora DSQL
- BlazingDB
- Chaos Mesh
- ClickHouse
- CouchDB
- Databricks
- DB9
- Dolt
- DSQL
- EraDB
- Fauna
- Firebolt
- FoundationDB
- Google Spanner
- HarperDB
- Iceberg
- kdb
- LanceDB
- Malloy
- MemSQL
- MonetDB
- Multigres
- NoisePage
- OpenDAL
- OxQL
- Pixeltable
- PostgreSQL
- QuasarDB
- RelationalAI
- RonDB
- ScyllaDB
- sled
- SpacetimeDB
- SplinterDB
- SQLite
- Striim
- Technical University of Munich
- TigerBeetle
- TonicDB
- Turso
- VillageSQL
- Weaviate
- YugabyteDB
- Alibaba
- ApertureDB
- Azure Cosmos DB
- Bodo
- Chroma
- CockroachDB
- CrateDB
- DataFusion
- dbt
- Doris
- DuckDB
- Exon
- FeatureBase
- Floe
- Gaia
- GoogleSQL
- Heron
- Impala
- Kinetica
- LeanStore
- MapD
- Microsoft SQL Server
- MongoDB
- MySQL
- Noria
- Oracle
- ParadeDB
- PlanetScale
- PRQL
- QuestDB
- RisingWave
- rqlite
- Sentry
- SLOG
- Spice.ai
- SQL Anywhere
- SQream
- Summingbird
- TeraData
- TileDB
- TopK
- Umbra
- Vitesse
- WiredTiger
Nov 18
2025
[Fall 2025] Optimizing the Table Scan Operator: I/O Minimization and Runtime Adaptivity
- Speaker:
- Benjamin Owad
- System:
- Snowflake
Table scan is a foundational operator in any analytical database and is often the primary bottleneck for a given query. This talk provides a technical deep dive into optimizations our team has developed for the table scan operator. First, we will discuss I/O reduction techniques, including pruning strategies to avoid reading unnecessary data and storage request coalescing to batch I/O... Read More
Nov 17
2025
[Future Data] Why Powering User Facing Applications on Iceberg is Hard
- Speaker:
- Benjamin Wagner
- System:
- Firebolt
- Video:
- YouTube
Firebolt is a Postgres compliant analytical database built for low-latency, high-concurrency analytics. These applications are usually powered by our fully managed storage and metadata layers. They support efficient caching and indexing, all while having multi-writer consistency. More recently, we’ve been investing heavily into our support for Apache Iceberg. Iceberg is not built to serve these types of low-latency applications. This... Read More
Nov 17
2025
Cortex AISQL: A Production SQL Engine for Unstructured Data
- Speaker:
- Anupam Datta
- System:
- Snowflake
Snowflake’s Cortex AISQL is a production SQL engine that integrates native semantic operations directly into SQL. This integration allows users to write declarative queries that combine relational operations with semantic reasoning, enabling them to query both structured and unstructured data effortlessly. However, making semantic operations efficient at production scale poses fundamental challenges. Semantic operations are more expensive than traditional SQL... Read More
Nov 11
2025
[Fall 2025] Open Data Infrastructure with Iceberg and dbt
- Speaker:
- Connor McArthur
- System:
- dbt
Apache Iceberg is now interoperable with most modern data platforms and compute systems. While Iceberg enables powerful new capabilities, real-world adoption still presents challenges for many organizations. In this talk, we will unpack Iceberg's architecture; demonstrate a novel architecture where multiple compute systems connect to the same underlying Iceberg catalog; and discuss the maturity and continued investment needed to ensure... Read More
Nov 10
2025
[Future Data] Mooncake: Real-Time Apache Iceberg Without Compromise
- Speaker:
- Cheng Chen
- System:
- Mooncake
- Video:
- YouTube
Apache Iceberg is great for large-scale analytics, but it was built for batch workloads. For streaming use cases, keeping tables fresh means writing snapshots more often, which creates excess small Parquet files, bloated metadata, and costly compaction that never ends. Updates and deletes make things worse because equality deletes push the burden to query engines, leaving readers slow and inefficient.... Read More
Nov 4
2025
Real Time Analytics Query Architecture Evolution @ Uber (Ankit Sultana)
- Speaker:
- Ankit Sultana
- System:
- Pinot
- Video:
- YouTube
We will talk about how Apache Pinot's query feature set has grown tremendously over the past few years and how that growth has shaped Uber's Real Time Analytics Query Architecture. We will dive into the different query engines in Apache Pinot and briefly discuss our legacy and unique Presto over Pinot architecture. Read More
Nov 3
2025
[Future Data] Multi-statement Transactions in the Databricks Lakehouse
- Speaker:
- Ryan Johnson
- System:
- Delta Lake
- Video:
- YouTube
The data lake architecture originally focused on self-standing tables in cloud storage, with catalogs as mere discovery aids. Modern lakehouse architectures add an ever-growing set of data warehousing capabilities to that original value proposition. Historically a key missing piece was multi-statement transactions -- Delta Lake supported single-statement single-table transactions, with ACID properties for changes made to that table. Sophisticated MERGE... Read More
Nov 3
2025
Transactions and Coordination in Aurora DSQL
- Speaker:
- Marc Brooker
- System:
- DSQL
Aurora DSQL is a new global, serverless, scalable relational database system, built at AWS. In this talk, I’ll dive into the architecture of DSQL, how it handles transactions, and how and why it was designed to minimize coordination. We’ll touch on transaction protocols, isolation, and virtualization. Read More
Oct 27
2025
[Future Data] Storage Metadata for Modern Cloud Databases
- Speaker:
- Joyo Victor
- System:
- SingleStore
- Video:
- YouTube
In modern database architecture, separating compute from storage unlocks powerful capabilities. Our tiered storage, “bottomless”, started by uploading files to remote object storage. This worked well until we wanted to create database branches pointing to the same remote storage. One branch does not know if it can delete a file that another branch depends on. To solve this, we built... Read More
Oct 21
2025
[Fall 2025] Astronomer / Apache AirFlow Tech Talk
- Speaker:
- Julian LaNeve
- System:
- AirFlow
Apache Airflow is the most popular data orchestration tool there is, downloaded over 40m times per month and used to power the data, ML, and AI platforms at OpenAI, Lyft, Airbnb, Uber, and Apple. At its core, Airflow allows you to define data workflows as DAGs using Python. We’ll do a deep dive on how Airflow came to be and... Read More