Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Sports tech's data pitch meets a harder question: who holds the keys

Sports analytics keeps getting sold as a data problem waiting for more AI, but the more useful debate in the dossier is architectural: whether agents should be handed raw data or a database to query.

SportAnalysisPeter LindqvistPublished: 27 September 20265 min readSources 4
Sports tech's data pitch meets a harder question: who holds the keys

Sports technology has spent years selling the same promise: more data, faster answers, smarter decisions. The market research headlines that cross the desk each week lean on growth rates and acronyms. The engineering reality is messier. Two documents in this week's stack are worth reading together, because they describe opposite ways of feeding an AI agent the numbers that sports organisations increasingly claim to run on.

Starburst published an analysis of two architectures for agentic data analysis on 11 September. It is written for enterprise data teams, not for clubs or leagues, but the example it uses is deliberately ordinary: a grocery store asking whether last week's egg promotion was profitable. Answering that means knowing whether the sale pulled in customers who would not otherwise have come, whether new customers returned, what else was bought in the same transaction, where those items sit in the store, and how profitable they are. That is a chain of follow-up questions across more than one dataset. A sports organisation faces the same shape of problem when it asks why a sponsorship campaign moved ticket sales.

Option 0: ship the data to the model

The first approach, which Starburst calls option 0, is to extract the relevant tables and send the raw data to the agent as a file or a direct pipe, letting the agent do its own processing. The company is blunt about the trade-offs. Sending terabytes of data to an agent can get prohibitively expensive when the agent charges per data item received or processed. Starburst says the performance disadvantages reduce the practicality of this option to the point where it is not a good one. That is why it carries the label option 0 rather than option 1.

For sports teams the arithmetic is unattractive anyway. A single match now generates tracking data, video-derived events, ticketing records, merchandise transactions and app telemetry, and the number of people asking questions about it is small. Paying per row to move that into a model is not a budget line most analytics departments can defend.

The second approach is the one Starburst recommends. Give the agent direct access to the database system and let it write its own SQL, iterating as it refines its focus. The database engine does what it is built for, processing local data with optimised query plans, while the agent supervises and sends successive or parallel requests until it has something worth returning. Starburst acknowledges the known downside: agents can be far more demanding than humans and can overwhelm a database with speculative queries while they work.

The database industry is laser-focused on supporting agentic workloads moving forward, and modern database systems are increasingly capable of handling this type of scalable data processing efficiently.

That is the argument in one sentence, and it is a sales argument as much as an engineering one. Starburst sells database systems. Even so, the point about where optimisation lives is hard to argue with: agents are improving, but they are not about to process terabytes of structured data as efficiently as engines built on decades of query research.

What the tooling actually looks like

Two open source projects in this week's stack show how the plumbing is being assembled at the smaller end. Adaca Analytics, published on GitHub on 8 September under the MIT licence, rebuilds the old Google Analytics dashboards on top of GA4 data. Daily rollups from the GA4 Data API or a BigQuery export land in a D1 database the operator owns, while realtime data stays live on Google. Six dashboards ship out of the box, every number opens into a detail page with its own trend and breakdowns, and the whole thing runs on Cloudflare Workers behind Cloudflare Access or built-in Basic Auth. There is no sign-in.

Plainoldanalytics, posted on GitHub on 13 September, takes the opposite route: bolt-on web analytics embedded inside a Go application. It is storage-agnostic, so importing the core package never pulls in a backend, and adapters exist for gin and chi. Traffic flushes to disk roughly once a second, and operators can set key and value properties per request, exclude routes from logging, or record literal route paths instead of path parameters.

Neither project is a sports product. Both are relevant because they show the default assumption shifting: the data stays where the operator put it, and the analytics layer is something you run rather than something you rent.

The privacy question nobody benchmarks

DataZen, a local-first database client posted to Hacker News on 29 August, states the problem directly: where does my data go? The tool ships under GPLv3, requires no account, and supports PostgreSQL, MySQL, SQLite and Redis by default, with optional drivers for MongoDB, ClickHouse, DuckDB and SQL Server. It includes a built-in MCP server so agents such as Claude, Cursor and Cline can query databases through the Model Context Protocol, and it can act as an MCP client too. Credentials are AES-256-GCM encrypted in the operating system keychain, and the project says only schema and query context are shared with the AI provider the user configures.

That last distinction is the one sports organisations should be asking about. Schema and query context is not the same as raw data, but it is not nothing either. A schema reveals what a club measures: which fan attributes it stores, how it segments ticket buyers, what it tracks about player workload. In a sport where injury data and contract negotiations sit in the same estate as ticketing, the metadata is commercially sensitive on its own.

The dossier does not tell us what any club or league has actually deployed, and the vendor material behind the market forecasts is not evidence of adoption. What it does show is a split. One camp argues that agents should drive the database and that modern engines can absorb the load. The other builds tools that keep the data local and send only context outward. Sports technology will inherit whichever answer its data teams choose, and the choice is being made now, in open source repositories rather than in press releases.

Comments 0

Sources

4
  1. 01An Analysis of Two Architectures for Agentic Data AnalysisEN
  2. 02We rebuilt the old Google Analytics on top of GA4's dataEN
  3. 03Show HN: Plainoldanalytics: Analytics Middlware for GoEN
  4. 04Show HN: DataZen – a local-first client for cross-database workflowsEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Peter Lindqvist

Peter Lindqvist

Sport, cars and travel

Peter Lindqvist covers sport, cars and travel for FLASH24, working from race results, manufacturer data and timetables rather than press releases. He checks entry lists and homologation papers against official series documents, and recalculates lap times, range figures and fare totals before anything goes out. He talks to team mechanics, rental desk staff and rail operators, and marks the Le Mans week and the winter timetable change in his calendar months ahead. Privately he drives an electric car, does his own garage repairs and plans rail routes across Europe, which is where most of his story tips start. He does not publish a number he cannot trace to a primary source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.