Aurora Query, a 16-Year-Old's Titan Breach, and the Kafka Migration
On 30 September, AWS made Aurora PostgreSQL able to query Apache Iceberg and Parquet data directly, one of a cluster of data-infrastructure changes landing on sports analytics teams this week.

Sports organisations sit on the same plumbing everyone else does: object stores full of Parquet, a warehouse, a transactional database, and a queue that keeps it all in sync. Three stories published in the last 72 hours describe that plumbing being pulled apart and reassembled, with consequences for how quickly a club, league or broadcaster can turn a match into a number.
Aurora adds DuckDB, and the ETL step disappears
AWS said on 30 September that Aurora PostgreSQL can now query operational data alongside data stored in Apache Iceberg and Parquet formats, using existing PostgreSQL applications and tools, without extract, transform and load pipelines or data duplication. Customers create PostgreSQL foreign tables that reference Iceberg or Parquet data in Amazon S3, Amazon S3 Tables or the AWS Glue Data Catalog. Aurora runs those queries through DuckDB's engine, which is embedded in PostgreSQL, according to the AWS What's New post. The capability is generally available on Aurora PostgreSQL 17.11 and 18.6 and higher in all AWS commercial and GovCloud (US) Regions, at no additional charge.
For a sports analytics stack, that is the difference between a nightly batch and a query that runs while the match is still going.
It also removes a known failure mode. "Accessing it has typically required pipelines that copy data from your data lake into Aurora, driving up costs and engineering work as schemas evolve," AWS wrote. AWS says external Iceberg REST Catalog-compatible catalogs federated through Glue Data Catalog work too. Latency-sensitive workloads can materialise Iceberg or Parquet data into native Aurora PostgreSQL tables using standard SQL, without an ETL pipeline.
Quantum and Kafka: the migration nobody wants to do twice
Object storage vendor Tigris described on 30 September why it moved asynchronous work out of FoundationDB, the only database it uses, and onto Kafka. The post is unusually candid about the trade-offs, and it maps closely onto sports data workloads where a tracking feed must not drop an event because a worker died.
Tigris had implemented the QuiCK paper, the queuing system Apple uses for CloudKit, on top of FoundationDB. It worked, but scheduling required many writes and scans, competing with user requests, and each task needed multiple writes to complete. New engineers had to learn custom code that had no standard implementation.
The company's objection to Kafka is worth quoting at length, because it is the objection most sports data teams will recognise. "Have you ever felt like a plastic bag drifting through the wind but unable to start again because of the sheer madness that comes with spending months permuting JVM flags to try to eke out a spectre's worth of performance so that your servers aren't constantly on fire?" the Tigris post asks. It then admits the migration was not a clean swap: queues remain in FoundationDB, while tasks such as garbage collection moved to Kafka. "We moved asynchronous tasks like garbage collection to Kafka, we can reduce the read and write load on FDB and shave off a good amount of that pesky custom code," the post states.
Anyone running a live sports data feed on a database-as-queue pattern should read it.
17.3 trillion rows, a 16-year-old, and a missing signature
The most-read data story of the week is not a product launch. The Register reported on 30 September that a 16-year-old security researcher named Faav found an authentication flaw in Microsoft's internal Titan analytics service, gained administrator access, and ran SQL with no valid credentials. The database he reached contained an estimated 17.3 trillion stored rows.
Titan is restricted to Microsoft employees through its web interface. Faav, working with an AI hackbot he built called Antares, reached Titan's API through an Azure Cloud Services host because Titan did not check the signature on a login token. Microsoft has since locked down the API and paid a $5,000 bug bounty. "It was 2 AM," Faav wrote in a blog about the findings. "I wanted to yell, or at least say something out loud, but my parents were asleep. So I just sat there staring at 17,333,335,124,315 and checked the math again."
The Register notes that Faav rewrote his blog post at Microsoft's request, cutting sections and numbers and rewording the impact before publication. Microsoft's statement to Faav, carried in the same piece, says: "Their submission and coordinated vulnerability disclosure helped us to better protect our customers by hardening our services."
For sports organisations, the relevant finding is not the row count. It is the lesson Faav drew: "Titan validated the contents of the JWT (tenant, audience, app ID, user) but never verified the signature, the most important part of any authentication check." Every analytics platform that ingests athlete, medical or fan data relies on the same token check, and the same failure would not announce itself.
Who owns the pipeline, and who gets to see inside it
Three other stories this week show that the governance layer around sports data is getting harder, not easier. The Guardian reported on 29 September that US Health Secretary Robert F Kennedy Jr laid out plans to connect medical and lifestyle data and search it with AI, citing Medicaid as "a really useful vehicle" with "hundreds of millions of lives in there". Sports medicine and wearable data sit adjacent to that pipeline, and the same questions about consent and secondary use apply.
The Guardian also reported on 30 September that more than 44,000 people have filed legal objections under article 21 of the UK GDPR to NHS England's Palantir-powered Federated Data Platform handling their personal information. Palantir's UK and Europe executive vice-president, Louis Mosley, has accused critics of "Palantir derangement syndrome". The company says its software has helped trusts record 117,000 additional operations and a 14.3% reduction in discharge delays for long-stay patients.
On the same day, 404 Media reported that the White House's Office of National Drug Control Policy uses the HIDTA grant programme to pull local license plate reader data from Flock, Axon and other vendors onto federal servers, in some cases forwarding it to the DEA's National License Plate Reader Program. Jeramie Scott of the Electronic Privacy Information Center told 404 Media: "If you're pissed about Flock then you should be pissed about this."
And NL Times reported on 30 September that most European data centres keep their environmental impact secret, with fewer than a quarter of larger Dutch facilities publishing electricity and drinking water figures, against a European Energy Efficiency Directive obligation in force for three years. Sport's data centres are subject to the same reporting gap.
What it adds up to
The technical direction is clear. Query engines are moving to where the data already sits, queues are being separated from databases, and the ETL copy that sports analytics teams have built for a decade is being deprecated by the vendors themselves. The governance direction is messier: more data linked, more objections filed, more of the underlying infrastructure undisclosed.
The two trends meet in the same place. A club that can query its lake from Aurora in seconds is also a club holding data that regulators, campaigners and attackers now treat as a target. The pipeline got faster this week. The arguments about it did not get simpler.
Sources
7- 01Aurora PostgreSQL now supports querying of Apache Iceberg and Parquet dataEN
- 02We used a database as a message queue. Now we use Kafka.EN
- 0316-year-old found Microsoft bug, got admin access to databases with 17.3 trillion rowsEN
- 04RFK Jr outlines expansive vision for collecting US health data at Maha eventEN
- 05More than 44,000 file legal objections to Palantir NHS platform handling their dataEN
- 06How Cities Are Forced to Funnel License Plate Data to a Massive Federal Surveillance ProgramEN
- 07Most data centers refusing to say how much water, electricity they useEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.