Lakebase Search: State-of-the-art full text and vector search for Postgres
Databricks introduces lakebase_vector and lakebase_text extensions for Lakebase Postgres, enabling state-of-the-art full-text and vector search directly within the database, eliminating the need for separate search engines.
Traditional OLTP systems struggle to meet the low-latency, high-accuracy search demands of AI agents, often requiring external search engines and complex ETL pipelines. Databricks addresses this by integrating a fast, scalable search engine directly into Lakebase Postgres through two new extensions: lakebase_vector for approximate neighbor search and lakebase_text for BM25 full-text search. These extensions are now generally available on AWS and Azure, offering a unified solution for both OLTP and search workloads.
The lakebase_vector extension significantly outperforms dedicated search engines in efficiency and scalability. On the VectorDBBench 100M benchmark, it delivers twice the throughput of the next best system while being four times cheaper than a cloud Postgres vendor using pgvector, even before autoscaling savings. It maintains high accuracy with a P99 latency of 71 milliseconds at 97% recall, ensuring reliable retrieval of nearest neighbors without sacrificing performance.
Lakebase Postgres now supports state-of-the-art search capabilities, enabling hybrid search with BM25 on over 100 million rows while reducing compute footprint by half compared to previous pgvector setups. Customers like Conexiom have adopted this approach, benefiting from a serverless database that scales to their needs without the complexity of separate search infrastructure.
The lakebase_vector extension resolves common pain points associated with pgvector, such as memory bottlenecks and slow ingestion. By decoupling storage from compute and leveraging quantized vectors, it ensures fast performance whether data is cached in RAM or cold on object storage. Index builds are parallelized and offloaded to distributed engines like Spark, reducing build times to minutes and enabling scalable, low-latency search operations.