Perplexity splits AI search storage across Pillar, Lorry and CobbleDB
Perplexity has detailed a production redesign that moves durable document state and update processing out of the query-serving path. It says CobbleDB cut batch-read latency by roughly fivefold and lowered estimated costs, though the results have not been independently reproduced.
Perplexity rebuilt its AI-search storage path around three components—Pillar, Lorry and CobbleDB—separating durable document state and batched updates from the system that serves search queries. The company says the resulting hot store reduced production batch-read latency by roughly fivefold compared with its previous DynamoDB-based path, but those figures come from Perplexity’s own measurements and have not been independently reproduced.
In Perplexity’s technical account of the redesign, CobbleDB is the read-optimized serving tier. Its keys are page identifiers derived from hashed URLs, while each value contains passages chunked in advance and vector embeddings for those chunks. The query path can therefore retrieve prepared records instead of processing documents at request time. For context on the company’s separate retrieval evaluation work, see its Q2D-Web benchmark for web-scale retrievers.
Pillar is the durable source of truth. Perplexity says it stores page metadata, chunks and embeddings in separate YTsaurus table families, retains multiple versions, and applies policies to determine which records should be exported to the NVMe-backed hot tier. When content changes or a record’s eligibility shifts, Pillar commits the durable state, consumed input and export intent together in a YTsaurus transaction. YTsaurus documents atomic visibility for changes made under its full transaction mode, although that documentation does not independently verify Perplexity’s configuration.
Lorry connects the durable layer to serving. The stateless consumer reads export records, groups them into files aligned with CobbleDB partitions, writes each batch to Amazon S3 under a unique identifier and registers it for ingestion. CobbleDB replicas then fetch and apply the batches independently and in chronological order.
The design moves transactional guarantees away from query serving. Perplexity says CobbleDB intentionally has neither transactions nor synchronized replicas, so a newly exported document may take a short time to become readable, and replicas can ingest at different rates. The durable update remains atomic in Pillar; the serving tier accepts delayed visibility to simplify read-heavy operation.
According to Perplexity, CobbleDB partitions data across nodes and keeps three replicas of each partition on different nodes. Each node uses RocksDB, serving cached data from memory and uncached data from local NVMe storage. A stateless router hashes page keys to partitions, requests them in parallel, prefers a replica in the same availability zone and can send a second request to another replica if the first is slow. Nodes use RocksDB MultiGet for batch reads; RocksDB’s documentation says the interface can process groups of keys more efficiently and parallelize I/O, but it does not validate Perplexity’s production results.
Perplexity reports that median batch-read latency fell from 31.4 milliseconds to 5.60 ms after the migration. It says p90 latency declined from 56.7 ms to 9.77 ms and p99 latency from 123 ms to 24.2 ms, improvements ranging from 5.08 to 5.80 times. However, the production comparison used live traffic at different times rather than identical requests in a controlled test. Perplexity describes batches of about 10 to 15 keys, average items of roughly 50 KB and traffic of approximately 200,000 requests per second in both systems. It also says load tests reached 500,000 requests per second without degradation. The company did not publish raw traces or a reproducible benchmark harness.
The company also estimates that CobbleDB costs at least 20% less than DynamoDB at every commitment level it evaluated. Perplexity says its internal model used storage size and read-and-write capacity estimates and excluded additional backup savings it expects from compression. The opened evidence does not include the full worksheet or an independent audit, leaving the comparison as an issuer estimate rather than a verified market price.
Perplexity says CobbleDB’s core contains about 40,000 lines of Rust and was built in two months by two human engineers working with hundreds of coding agents. It plans to release the project as open source, but its post identifies no release date, license or public repository.
More news

xAI launches Team Bots for shared workflows

AWS adds xAI’s Grok 4.7 to Amazon Bedrock

OpenAI adds $5 million and up to $5 million in credits to Lenfest AI program
