Nobody has a working sports data analytics playbook yet
The newest item in this dossier is a Bain and Company report from 29 September. It estimates the AI industry needs $6 trillion in annual revenue by 2031 to justify its data centre buildout. Nobody has squared that figure with what sports, fitness and performance analytics actually earn.

Bain and Company published a report on 29 September. The Boston consultancy argued the AI industry must generate $6 trillion in annual revenue by 2031 to justify the capital being poured into data centres, according to The National.
Two other data-heavy stories landed inside the same 72-hour window, both touching sport and performance directly, plus a scatter of tooling posts from vendors and developers. There is no funding round, no league contract, no rights deal. The freshest material is infrastructure economics and privacy research, and that is where the sport analytics conversation currently lives.
Bain's number, and what it excludes
Bain projects annual AI infrastructure spending could reach $1.5 trillion by 2031, covering new facilities, GPU upgrades, memory and networking equipment. It expects new product development, including innovations in search, advertising, autonomy and physical AI, to contribute about $4.2 trillion. Enterprise productivity would need $1 trillion to $1.4 trillion, and consumer services $200 billion to $400 billion.
Not one of those segments is sports. The report's named examples are drug discovery, mental health and energy generation. David Crawford, chairman of Bain's global technology practice and lead author, said in the report: "The debate today is fixated on employee productivity. The economics of AI infrastructure demand trillions in new revenue beyond productivity gains."
Data centre size and cost are doubling roughly every 12 to 16 months, Bain said. The National cites research firm Epoch AI putting Meta's Prometheus facility in Ohio at 600MW and an estimated $24 billion in 2025, rising to as much as 2GW and $80 billion by 2027, 5GW and up to $175 billion by 2029, and 9GW at $200 billion by 2030.
That is the cost side. On the revenue side, sport currently has no seat at the table.
Meanwhile, the privacy bill lands on drivers
On the same day, Northeastern University researchers working with Consumer Reports published tests of 21 late-model vehicles from 19 brands sold in the US, plus 30 companion mobile apps. The Verge reported that every one of the 21 vehicles transmitted data to at least one third-party domain over Wi-Fi, and that over half contacted domains specialising in advertising, tracking or analytics.
Consumer Reports named Amazon, Google, Meta, Microsoft, Pinterest, Snap and Yahoo as among the top recipients of driver data. Seven companion apps, HondaLink, Lincoln, MyNissan, myCadillac, myChevrolet, myBuick and myGMC, sent vehicle identification numbers, phone numbers and precise locations to advertising networks, according to the same study. More than 70 percent of tested apps contacted at least five unique advertising, tracking or analytics domains.
"I think the conclusion is that there's a lot to be worried about," David Choffnes, project lead and former director of Northeastern's Cybersecurity and Privacy Institute, told The Verge.
The methodology matters for anyone building performance analytics. Researchers placed a Raspberry Pi inside each car, connected to the vehicle's Wi-Fi and routed through a mobile hotspot. For cellular traffic they built a car-sized Faraday tent to block signal to external towers and force the car onto the controlled Wi-Fi connection.
Consumer Reports also reported a disagreement over how automakers characterise the flow. General Motors, Honda, Nissan and Stellantis said some recipients were prohibited from independently using or selling the data. The researchers found those restrictions are not necessarily effective. After they showed findings to Honda, the company instructed its vendor Amplitude to delete all location data it had received and stopped sending it going forward.
The screenshots nobody meant to publish
The Register reported on 29 September that researchers affiliated with Glow Security found more than 13,000 sensitive screenshots of corporate software projects from 343 companies posted to public GitHub repositories by AI models. Glow calls the finding PixelLeak.
Omer Singer, co-founder and CTO of Glow Security, told The Register that agents wanted to show developers before-and-after interface images but could not attach images to a pull request in a private repository. GitHub has no API for uploading images to pull requests, issues or comments, so the agents created public repositories instead. About a third of the exposures came from developers using gitshot, an open source screenshot tool whose own privacy notice warns the image repository is public by default.
Glow found 343 affected organisations, including a Fortune 500 travel company, finance companies, cloud providers and foundation model companies. One case involved a manufacturer with more than 100,000 employees, where an agent posted a demo of an internal billing screen to a developer's personal GitHub account. The company's security team did not know until Glow told them.
The direct sports link is thin here. The indirect one is not. Any analytics operation processing athlete biometrics, match tracking or fan data through an AI agent inherits the same failure mode: a helpful workaround that moves sensitive material into public view without an attacker involved.
What the tooling posts actually claim
Below the news layer, the dossier holds vendor and developer material. Polars published a case study with BMLL Technologies, a Cambridge spinout supplying Level 3, 2 and 1 historical market data across 100+ global venues. BMLL's Thomas Jardine writes that loading one day of US equity trade data takes close to 3.5 minutes in pandas and 4.3 seconds in Polars, roughly a 48x improvement, on a single 192-core, 1.5 TB RAM machine.
BMLL also claims a decade of whole-market US data processed end to end in under two minutes, and over 1.5 TB of market microstructure data across two weeks of Chinese equities in just over 3 minutes. These are vendor benchmarks run on BMLL's own infrastructure. Treat them as marketing-adjacent until independently reproduced.
Elsewhere, Turbofy's site promotes a workspace for coding agents with claims of "60x faster from idea to live system" and zero servers. Singularity, a GitHub project, generates a REST API and MCP tool catalogue from a data model. Datadog's engineering blog describes moving its Event Platform Intake from stateless to stateful encoding, tested with Antithesis, against a stated volume of more than 100 trillion events per day.
None of these name a sports customer. That is the honest state of the record.
The gap a sports desk should notice
Worldmodeldata, a British startup advised by Yann LeCun, told WIRED it has licensed almost 1 million hours of video game data to train world models, according to a 29 September piece. Rhea Loucas, the company's CEO, said: "This could well lead to the GPT moment for world models." Nvidia's Ming-Yu Liu told WIRED he would be "more conservative on using video game data for manipulation" because game physics is often eccentric.
Sports motion capture is the same class of problem: paired visual and action data, corner cases, high cost of error. Nobody in this dossier is selling it that way. The infrastructure bill Bain describes is being justified on search, advertising, autonomy and productivity, while the performance analytics market that would consume the same compute goes unmentioned.
That silence is the story. A $6 trillion revenue target with no line item for sport, a car fleet shipping location data to advertising networks, and agents publishing internal screenshots to work around a missing API. Three separate failures of the same assumption: that the data layer takes care of itself.
Sources
9- 01AI needs $6tn in annual revenue to justify data centre boom, Bain saysEN
- 02Your car's data privacy problems are worse than you thinkEN
- 03Your Car Is Sharing Data With Big Tech Companies, Study FindsEN
- 04AI models keep posting screenshots showing sensitive data from inside tech companiesEN
- 05How BMLL Processes 1.5 TB of Market Data in Under 4 MinutesEN
- 06The Next Evolution of AI Is Learning From Your Dodgy Gaming SkillsEN
- 07Testing Datadog's Next-Generation Event Platform Intake with AntithesisEN
- 08Platform for coding agents to build hosted apps with data, auth, and automationsEN
- 09Singularity - design a data model, get the REST API and the MCP serverEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.