Postgres, Luxir and SlopShape: the search stack is being rebuilt while publisher traffic falls
PlanetScale published a deep technical explainer on Postgres full-text search on 27 September, the newest item in a week of search infrastructure releases that lands as publishers report steep falls in search referrals.

The newest development in the search stack this week did not come from a search engine at all. On 27 September, PlanetScale published "Anatomy of a (Postgres) Search Engine," a technical explainer on how inverted indexes work inside Postgres and how they differ from the b-trees most developers reach for first.
PlanetScale's writeup is a reminder that search is now a database feature, not just a product category. "A b-tree is useless for matching content in the middle of a string," the company writes, before showing the pattern it wants developers to stop using: a query that checks every row in a table with a leading-wildcard LIKE. The alternative is an inverted index keyed by each individual word in a field. According to the post, that is how "essentially all full-text search engines" locate documents. It is a small architectural point with large consequences for how quickly a database can answer a question.
The explainer is unusually concrete about the machinery. An inverted index has a term dictionary and postings lists, and it may add positional and frequency data on top. Postings lists are sorted, then compressed. PlanetScale gives the example of an index with 300,000 documents, 100,000 of which contain the word "who." Storing each ID literally would take 19 bits per posting, but the average gap between successive IDs is three, which fits in two bits. Dense lists can go further and store a bitmap, which becomes optimal once around half of all documents contain a term. That is a compression detail, but it explains why search over large corpora is affordable at all, and why the query engine can take the union or intersection of two postings lists in a single O(n) pass. PlanetScale also notes that vector instructions on recent CPUs can OR or AND 128, 256, or even 512 bits in one instruction when lists are stored as bitmaps.
Luxir pitches full-spectrum search as one engine
Five days earlier, on 22 September, a separate project called Luxir published its own pitch: an open source engine that combines full-text, vector and faceted search with analytics in one request. Luxir's site frames the argument in cost terms rather than feature terms. The engine is "designed for the most search per core, per gigabyte, and per dollar," it says, and cloud customers "pay for inefficiency forever."
The architecture claims are specific. A work-stealing scheduler shares indexing, merging and query execution across cores. Network IO is asynchronous, so a slow client does not hold a core, and segments are immutable and memory-mapped, so they are read from the page cache. Postings decoding, scoring and vector distance run on SIMD paths, with block-max pruning for top-k requests.
Luxir also takes a position on counts that matters to anyone building a search UI. Facet counts are "always exact by default," the project says, and total result counts are exact when the request asks for them and pruned when it does not. Queries, facets and metrics run in the same pass over the same index view, so one round trip returns documents, sidebar counts and header numbers together.
Both posts describe plumbing most users never see. But they are the layer where the economics of search get decided, and the economics are changing fast.
A detector that reads structure, not vocabulary
On the same day as Luxir's launch, a paper on arXiv took aim at the content that fills these indexes. "SlopShape: Identifying AI-Generated Commercial Web Content," submitted on 14 September and revised on 17 September, asks whether AI-generated text can be identified from structural signatures, meaning how information is presented, in what order, with what evidence, and in what voice.
The author, Jochen Madler of Sitefire, replicated the StoryScope work by Russell et al. on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains, against 11,250 AI mirrors from five frontier models. A 214-feature instrument, applied by an LLM and validated in a human gold-annotation session with human-human kappa of 0.928 and human-model kappa of 0.946, detected AI posts from its 187 structural features alone at 98.0 macro-F1 on held-out companies. Rewording every AI post with its own model left the score unchanged at 98.1.
The paper also claims attribution. AI posts share "a tidy, self-announcing shape," 79.3 percent are attributed to the correct source model against a 16.7 percent chance rate, and human posts occupy rare structural configurations. The pipeline, instrument, prompts and code are released.
That result matters for search because structure is exactly what ranking systems and AI answer engines consume. If commercial web content has a detectable shape, then the corpus that search engines index is becoming more uniform at the same time as the tools for indexing it get faster and cheaper.
Publishers are the ones paying for it
The surrounding headlines in the dossier point the other way. Recent coverage cited in the dossier includes a report that Google search decline steepened to 40 percent year over year for publishers, dated 24 September, and a separate item saying Google Search Profile Badges expose the publisher traffic crisis, dated 16 September. Older items describe search engine traffic down across the web, with some small websites seeing a 60 percent drop, and Google AI search threatening small publishers as conversational AI cuts website traffic. Some of those dates sit outside the last week, so they are background rather than news. But they set the frame for the technical releases above: the cost of building a search index keeps falling, and the traffic that index sends back to publishers keeps falling with it.
Two other dossier items illustrate how wide the field has become. On 26 September, CleanTechnica opened submissions for its own book publishing arm, offering authors a route to publication for an upfront fee of $3,000 to $5,000 plus a percent of each book sold, with a Google Hangout scheduled for 30 September. On 24 September, Wiley's engineering content hub published a white paper, sponsored by Hendrix by Marmon Utility, on substation exit construction. The paper argues that reliability in the first spans out of a station carries unusual weight, because a single contact-driven fault there can interrupt many circuits at once.
Neither is a search story on its face. But both are commercial content aimed at a specific audience, produced under a sponsor's name or a publisher's brand, which is the same category the SlopShape paper tries to classify. The paper's claim is that such content has a shape: a particular order, a particular kind of evidence, a particular voice. If that holds, the next generation of search engines will have to decide what to do with it.
For now, the infrastructure side is moving faster than the policy side. PlanetScale is shipping TIN, its full-text search index for Postgres, and says a deeper article on its features, performance and implementation is coming. Luxir is downloadable, with an HTTP/JSON API on port 9400 and a data directory for persistence. The arXiv paper is public, with code and artifacts attached.
What none of them settles is where the readers go. Search engines are getting better at finding documents, and publishers are reporting that fewer people arrive at them. That gap is the story the tooling has not answered yet.
Sources
5- 01Anatomy of a (Postgres) Search EngineEN
- 02Luxir: Open-source hybrid search engineEN
- 03SlopShape: Identifying AI-Generated Commercial Web ContentEN
- 04Publish Your Book Through CleanTechnica PressEN
- 05Engineering the Substation Exit for Reliability, Capacity, and ExpansionEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.