Databricks launches NEAREST BY Join for batch vector search
Databricks has turned exact and approximate batch vector search into a native SQL join, using Photon execution and an optional Delta-based IVF index.
Databricks introduced NEAREST BY Join on October 5, making exact or approximate batch vector search a native SQL operation in Databricks Runtime. Data teams can express the batch as one relational query over data already held in the platform.
The company presents that approach as an alternative to moving embeddings into a separate vector store and maintaining a synchronization pipeline for offline work such as deduplication, entity resolution, tagging and recommendation jobs. The feature provides a batch-processing path inside the runtime, distinct from Databricks vector and full-text retrieval in Lakebase Postgres.
For each row on the left side of the join, NEAREST BY returns up to k rows from the right side, ordered by a distance or similarity expression using values from both sides. The left side drives the search; the right side is searched. The documented syntax supports INNER and LEFT OUTER joins, with the latter retaining query rows when no candidate is found. Users can select EXACT for exact top results or opt into APPROX so the optimizer may use an approximate strategy.
Databricks says the optimizer lowers the operation to a plan with a cross join, pairwise scoring and grouped top-k, allowing the job to use distributed execution, spilling and retries. In Photon, the company says native SIMD vector functions and a fused operator can replace the cross join, scoring projection and partial top-k with a blocked matrix-multiplication kernel. Databricks claims this setup reuses buffered vectors, avoids writing the full score matrix to memory and limits emitted candidates when top-k is fused. Those efficiency claims have not been independently benchmarked.
For approximate searches, Databricks says the optimizer can use an inverted-file, or IVF, index stored as an ordinary Delta table and liquid-clustered by centroid ID. Files tied to irrelevant centroids can then be pruned before reading. The company also says index refreshes are transactional and incremental. Rows added since the last refresh are searched through a brute-force compensation branch, so stale indexing reduces pruning rather than search quality.
In company-reported evaluations, base tables ranged from 100,000 to 5 billion vectors, query batches reached 10 million vectors and indexed tests targeted at least 96% recall@K. Databricks said an indexed approximate search completed a 10-million-row deduplication self-join in minutes and searched 10 million queries against 100,000 reference vectors in under a minute. It also reported minute-scale runs for 1 million queries against sets of 10 million and 1 billion vectors.
The results have not been independently reproduced. The announcement also omitted several details needed for reproduction, including workload-specific cluster configurations, embedding dimensions, k values, datasets, precise elapsed times and costs.
Databricks documentation says NEAREST BY applies to Databricks SQL and Databricks Runtime 18 LTS and above. It documents k values from 1 to 100,000, requires matching ARRAY dimensions for vector scoring functions and says the clause is not supported on streaming DataFrames or Datasets. The opened sources do not specify whether the feature is generally available or in preview, or whether cloud and region restrictions apply.
More news

Databricks adds Iceberg catalog handoff APIs in private preview

Databricks adds branch-based restores to Lakebase Postgres

Databricks makes native IP Functions generally available
