Building for the AI Era: Lakebase, Streaming, and Lakehouse Innovations at VLDB 2026
Databricks will present innovations at VLDB 2026, including Lakebase, streaming improvements, and automated optimizations, addressing AI agent workload demands and real-time analytics.
Databricks will present four accepted papers at VLDB 2026, including work on Lakebase, Spark Structured Streaming, and automated Lakehouse optimizations. Co-founder Reynold Xin will deliver the opening keynote, reflecting on the evolution of database technology from Shark to Lakehouse and introducing new paradigms like Lakebase and LTAP to support AI agent workloads. Lakebase decouples PostgreSQL compute from storage, enabling sub-second cold starts and Git-like database workflows while supporting low-latency analytics on live transactional data.
Spark Structured Streaming, used in millions of weekly jobs, has evolved with microbatch pipelining improving throughput by up to three times. New stateful APIs simplify complex business logic, and fine-grained access control has been added. The system now supports clustering tables by keys to enhance query performance, though manual key selection does not scale across millions of tables.
AutoLiquid automates clustering key selection using heuristics and shadow verification, outperforming customer-selected keys in over 95% of evaluated workloads. Ultron, a history-based query optimization framework, leverages repetitive analytical workloads to improve optimizer choices, reducing median join latency by 25% in production workloads.
Databricks will also demonstrate Enzyme, an incremental view maintenance engine, and LakehouseRT, designed for real-time low-latency analytics over open lake storage. The combination of Lakebase and LakehouseRT aims to deliver the first true Lake Transactional Analytical Processing system, with details available at the Databricks booth during VLDB 2026.