---
title: "Perplexity Launches Photon, a Rust-Built Retrieval Engine That Slashes Search Latency by 92%"
summary: "Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust, replacing a forked open-source engine. The company reports that production p99 latency dropped from roughly 800 ms to 65 ms. Photon now handles all production search traffic and powers a new Fast Search mode in the Perplexity Search API, priced at $1 per 1,000 requests."
url: "https://theauthority.report/articles/perplexity-launches-photon-a-rust-built-retrieval-engine-that-slashes-search-latency-by-92-b62b25c5"
publisher: "The Authority Report"
section: "AI Discovery"
author: "The Authority Desk"
datePublished: "2026-10-04T16:50:20.762Z"
dateModified: "2026-10-04T16:50:20.762Z"
sources:
  - title: "Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms"
    url: "https://www.marktechpost.com/2026/09/30/perplexity-introduces-photon-a-rust-based-retrieval-engine-that-cuts-p99-latency-from-800-ms-to-65-ms/"
---

# Perplexity Launches Photon, a Rust-Built Retrieval Engine That Slashes Search Latency by 92%

*Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust, replacing a forked open-source engine. The company reports that production p99 latency dropped from roughly 800 ms to 65 ms. Photon now handles all production search traffic and powers a new Fast Search mode in the Perplexity Search API, priced at $1 per 1,000 requests.*

Perplexity has released Photon, an in-house retrieval and ranking engine written in Rust, replacing an open-source engine the company had previously forked for its AI-native search stack. According to Perplexity, Photon now handles retrieval and ranking for all production traffic and also powers a new Fast Search mode in the Perplexity Search API.

## Why Perplexity Rebuilt From Scratch

Perplexity said the previous engine hit three scaling limits as its index grew. At the tail, production p99 latency sat near 800 ms; because the dataset exceeded available RAM, cold reads triggered page faults that stalled queries. During disk index merges, p99 climbed to roughly 1.2 seconds for stretches of 10 to 15 minutes. Recovery by deploying and syncing an extra cluster could take more than a week. The company said it concluded that building from scratch was simpler and cheaper than continuing to maintain its fork.

## How Photon Works

Photon uses a broker-and-shard architecture: a load balancer routes each request to a broker, which fans out to a shard group, watches for timeouts, merges candidates, and fetches key document fields. Each shard independently handles retrieval, initial ranking, and second-stage ranking.

Key engineering choices include adaptive posting lists, where short lists sit inline within a single page and longer lists split into fixed document-ID-range blocks; a budgeted WAND-like traversal algorithm that checks cheap presence scores before reading exact term frequencies; compact per-document "Docblob" records using Elias-Fano encoding; and batched asynchronous disk reads through io\_uring with a lock-free cache using CLOCK eviction. Indexers build versioned shard indexes from YTsaurus tables on dedicated nodes, and a controller warms caches with replayed search-log queries. Perplexity reports a full web index now builds in a single-digit number of hours.

## Production Results

Perplexity reports that p99 retrieval and ranking latency fell from about 800 ms to about 65 ms, covering Photon's stages only. The engine runs on about 20% fewer serving machines than the previous content nodes and stores about 2.5 times as much data per document, which Perplexity says it used to improve ranking quality. The company also notes that pinning the same dataset with mlock would require an estimated 4.6 times the resident memory Photon currently uses, and that index version switches no longer cause latency spikes.

## Fast Search Mode and Benchmark Results

Fast Search pairs Photon with lighter ranking tuned for agentic workflows and is available via the Perplexity Search API at $1 per 1,000 requests by setting search\_type: "fast" on a POST request to /search. Perplexity reports single-call latency of 160 ms at p50 and 230 ms at p95 for Fast Search.

Perplexity tested Fast Search against its default preset across six benchmarks — WideSearch, BrowseComp, DSQA, FRAMES, SEAL-0, and SEAL-Hard — covering 3,554 tasks. Fast Search scored 64.3% at an estimated model-plus-search cost of $59.73, compared with 64.0% at $187.60 for the default preset, a cost reduction the company characterises as approximately 68%.

Perplexity also disclosed trade-offs: on internal long-tail benchmarks, relevance as measured by DCG fell from 2.45 to 2.21, and answer availability dropped from 0.596 to 0.567, a loss of 2.9 percentage points. The company recommends Fast Search for routine agent loops and the default preset for hard or ambiguous queries.

Photon is not open source and cannot be self-hosted; it is available only as a hosted API.

## Sources

- [Perplexity Introduces Photon: A Rust-Based Retrieval Engine That Cuts p99 Latency From 800 ms to 65 ms](https://www.marktechpost.com/2026/09/30/perplexity-introduces-photon-a-rust-based-retrieval-engine-that-cuts-p99-latency-from-800-ms-to-65-ms/)

*Cite as: The Authority Report, "Perplexity Launches Photon, a Rust-Built Retrieval Engine That Slashes Search Latency by 92%", https://theauthority.report/articles/perplexity-launches-photon-a-rust-built-retrieval-engine-that-slashes-search-latency-by-92-b62b25c5 (as of 2026-10-04).*
