Next upSF Pitch Night by the AI Collective - #SFTechWeek
News

Databricks launches Lakebase Search for vector and full-text retrieval in Postgres

Databricks has made Lakebase Search generally available on AWS and Azure, bringing vector, keyword and hybrid retrieval into its Lakebase Postgres service.

D
Sep 30, 2026 · 2 min read

Databricks announced general availability of Lakebase Search on AWS and Azure, bringing vector, keyword and hybrid retrieval to its Lakebase Postgres service. Developers can now search operational data in the same database rather than maintain a separate search system and data pipeline.

The launch post was published on September 28, 2026, but the company’s Lakebase release notes list September 18 as the general availability date. Together, the records establish a September GA release without resolving the first day customers could use the feature.

Lakebase Search is delivered through two Postgres extensions. lakebase_vector handles approximate nearest-neighbor retrieval through the lakebase_ann index type, using pgvector-compatible vector types, distance operators and query syntax. lakebase_text adds BM25 keyword ranking through a lakebase_bm25 index over standard PostgreSQL tsvector data. BM25 ranks documents by term frequency, document length and how rare a term is across the collection.

For vector search, Databricks said the system organizes vectors into hierarchical IVF clusters stored as contiguous blocks, then applies RaBitQ binary quantization to compress each vector to roughly one bit per dimension. A query uses those compact representations to find likely matches, reads the promising blocks and reranks a shortlist against full-precision vectors. Keeping durable index data separate from compute allows nodes to cache the working set and suspend when idle.

For text retrieval, Block-Max WAND top-K pushdown skips blocks that cannot contain a high-ranking result instead of scoring every match. The Azure Databricks documentation shows hybrid search running vector and keyword queries separately, then combining the rankings with Reciprocal Rank Fusion in SQL.

According to the Azure documentation, Lakebase Search requires PostgreSQL 16 or later. Enabling it restarts every compute resource in a project, drops active connections and cannot be reversed. Users then install the vector and text extensions in each database. The release extends the Lakebase product line beyond developer workflows such as branch-based database development with Consort.

Databricks also published performance results that have not been independently reproduced. In its VectorDBBench test using 100 million LAION vectors, the company said lakebase_vector delivered twice the throughput of the next-best tested system and cost one-quarter as much as a cloud Postgres service running pgvector. Databricks reported 71-millisecond P99 latency at 97% recall, while noting that pgvector and DiskANN were each tested on only one large instance. The post does not include enough configuration and pricing detail to normalize the comparisons across all tested systems.

The company also said Lakebase could serve 100 million 768-dimensional vectors on one Lakebase Compute Unit and resume a first query after scaling to zero with a measured P90 latency of 1.13 seconds. Databricks said customer Conexiom ran hybrid BM25 search across more than 100 million rows using half the compute footprint of its previous pgvector setup, with threefold lower database spending and fivefold higher throughput. Those customer figures have not been independently verified.

More news