Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Search Engine Traffic and Publishers: What the Newest Open Source Search Tools Reveal

On 27 September, PlanetScale published a technical deep dive into how full-text search engines work. The machinery that decides which pages publishers get traffic from is being rebuilt from the index up. The post landed the same day two smaller search projects were still drawing attention on Hacker News. One of them had been submitted just 7.5 hours earlier.

Media & internetExplainerGrace OkonkwoPublished: 28 September 20264 min readSources 6
Search Engine Traffic and Publishers: What the Newest Open Source Search Tools Reveal

On 27 September, PlanetScale published a long technical explainer titled "Anatomy of a (Postgres) Search Engine." It is not about publishers. It is about inverted indexes, term dictionaries and postings lists, the data structures that let a search engine find a document by the words inside it rather than by scanning every row.

The timing matters. Google search referrals keep sliding for many publishers, and who controls the index has become a media business question, not just an engineering one. The PlanetScale post walks through how a full-text index maps each term to the documents containing it. The company says a standard b-tree index can quickly find a name, a price range or a creation date, but it is "useless for matching content in the middle of a string." That is the job of an inverted index, which stores a term dictionary and postings lists, often compressed with gap encoding or bitmaps. PlanetScale is using the piece to introduce TIN, its full-text search index for Postgres. A deeper look is promised in a separate article.

The search stack is opening up

Two other projects in the dossier point the same way. Luxir, an open source hybrid search engine, describes itself on its homepage as built for "the most search per core, per gigabyte, and per dollar." It combines full-text, vector and faceted search in one engine, with an HTTP/JSON API on port 9400, BM25 ranking with block-max pruning, and exact facet counts. Luxir's page was published on 22 September, six days before this article.

Both PlanetScale and Luxir are selling efficiency. PlanetScale's post notes that postings lists are "almost always sorted, then compressed," and that dense lists can be stored as bitmaps where a single bit marks whether a document contains a term. Luxir makes a similar pitch: immutable, memory-mapped segments, SIMD decoding, a work-stealing scheduler, and no garbage collector. "In the cloud, you pay for inefficiency forever," Luxir's homepage says. That is a vendor claim, not an independent benchmark, but it is the same argument.

"A b-tree is useless for matching content in the middle of a string. That's where full-text search indexes come in." PlanetScale, 27 September

The publisher angle is indirect, and it is worth being precise about it. None of these sources contains a number for Google referrals, or for any publisher's traffic. What they show is that the building blocks of search, inverted indexes, vector ranking, faceting, are increasingly available as open source or as add-ons to general-purpose databases. For a publisher, that does not restore lost search traffic. It does mean the cost of running your own site search, or a niche vertical search product, keeps falling.

Older signals, still the backdrop

The dossier's other recent items are further from the media beat. A Wiley white paper published on 24 September covers substation exit construction and aerial cable systems, sponsored by Hendrix by Marmon Utility. An arXiv paper submitted on 14 September and revised on 17 September, "SlopShape," describes a 214-feature instrument that identifies AI-generated commercial web content from structural signatures alone, reporting 98.0 macro-F1 on held-out companies against 11,250 AI mirrors from five frontier models. A GitHub repository from 22 September publishes optimized runtimes for three vision-language-action models on NVIDIA Jetson AGX Thor, with latency figures for ABC-VLA, MolmoAct2 and pi0.5.

The SlopShape paper is the one with an obvious adjacency to publishing. Its author, Jochen Madler of Sitefire, writes that word-level AI detectors are brittle under rewording, and that structural features hold up better: 98.1 macro-F1 when every AI post is reworded by its own model. The paper says AI posts "share a tidy, self-announcing shape," and that human posts "occupy rare structural configurations." It attributes 79.3% of AI posts to the correct source model against a 16.7% chance rate. The paper is a preprint and the author discloses an affiliation with a commercial vendor, Sitefire.

Read together, these sources describe a search and content ecosystem under active reconstruction. PlanetScale is extending Postgres downward into search. Luxir is offering an open source alternative built for density. SlopShape is trying to measure the AI content that fills the resulting indexes. And the Wiley white paper is a reminder that most of the dossier is not about media at all, which is itself a signal about where the technical energy is going.

The honest gap: the dossier contains no fresh measurement of publisher search traffic, no Google statement, and no named publisher reporting a decline. The most recent datable development is PlanetScale's 27 September post. Anything about referral losses is context from outside this dossier, and should be treated as such.

Comments 0

Sources

6
  1. 01Anatomy of a (Postgres) Search EngineEN
  2. 02Luxir: Open-source hybrid search engineEN
  3. 03SlopShape: Identifying AI-Generated Commercial Web ContentEN
  4. 04Engineering the Substation Exit for Reliability, Capacity, and ExpansionEN
  5. 05Show HN: Optimized runtimes for three VLAs on Jetson ThorEN
  6. 06Show HN: ParkourNoteEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Grace Okonkwo

Grace Okonkwo

AI, models and technology

Grace Okonkwo covers AI, models and technology for FLASH24, working from primary sources such as model cards, API documentation and benchmark papers rather than vendor summaries. She checks training data provenance, evaluation conditions and reported scores against the underlying datasets before any figure reaches print. She interviews researchers and engineers directly, tracks release calendars from major labs, and compares successive model versions on the same tests. Her own self-hosting, home-network and documentation-reading habits feed straight into that desk, since she tests tools on her own hardware first. She does not publish benchmark claims without a reproducible method.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.