By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Perplexity Launches Rust Retrieval Engine Photon
Perplexity has launched Photon, an in-house retrieval and ranking engine developed using the Rust programming language. This new engine replaces a previous open-source engine that Perplexity had forked to support its AI-native search stack. Photon is now responsible for handling all retrieval and ranking tasks for Perplexity's production traffic. Additionally, it powers a newly introduced "Fast Search" mode available through the Perplexity Search API. Perplexity reports that Photon achieves single-call latencies of 160 milliseconds at the 50th percentile (p50) and 230 milliseconds at the 95th percentile (p95). The "Fast Search" mode can be accessed by setting the search_type parameter to "fast" in a POST request to the /search endpoint, with a cost of $1 per 1,000 requests. However, Photon itself is not open-source, meaning it cannot be self-hosted by users. The decision to build Photon from scratch stemmed from limitations encountered with the previous engine as Perplexity's index grew. These limitations included tail latency issues, where production p99 latency approached 800 milliseconds. The dataset's size exceeded available RAM, preventing the use of memory locking (mlock), and cold reads frequently triggered significant page faults that stalled queries. Furthermore, the old engine experienced "merge spikes" during disk index fusion, causing p99 latency to rise to approximately 1.2 seconds for durations of 10 to 15 minutes. The recovery process for the previous system was also protracted; deploying and syncing an additional cluster could take over a week, and this recovery period often led to an increase in partial responses. The Perplexity team concluded that developing a new engine internally was both simpler and more cost-effective than continuing to maintain their modified fork of the open-source solution. Photon operates by routing incoming requests to a Photon broker via a load balancer. This broker then fans out the request to a designated shard group and monitors for any timeouts. Within each shard, the processes of retrieval, initial ranking, and second-stage ranking are executed. Subsequently, the broker consolidates the candidate results and retrieves essential document fields. The engine employs adaptive posting lists, where shorter lists are embedded directly within a single page, while longer lists are segmented into blocks based on fixed document ID ranges. Sparse blocks utilize sorted offset arrays and implement a galloping search algorithm, whereas dense blocks leverage bitmaps for efficient membership checks via single bit lookups. Photon also utilizes budgeted traversal, a WAND-like algorithm that divides lists into driving and probe lists. This approach prioritizes cheap presence checks to establish a candidate's maximum potential score before proceeding to read exact term frequencies, which is only done if a candidate can surpass the established threshold. Document records, referred to as "docblobs," store compact frequency and field information for each document.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.