Akka tests AI porting across 65 open source projects using 9.4B tokens
On October 5, Akka reported that a spec-driven AI workflow ported 65 open source projects, consuming 9.41 billion tokens in 99.3 hours while improving code or performance in 57 of them.

The experiment ran for 99.3 hours. It consumed 9.41 billion tokens. Akka used this compute to test how well AI agents could port open source software when guided by structured specifications.
Spec-driven porting at scale
According to InfoQ, Akka analyzed all 65 projects in an initial tranche, generating specifications and implementing up to 10% of each project's surface area. They then selected 10 projects for complete implementation based on system characteristics. The delivery harness cycled through discovery, specification, porting, benchmarking, and improvement. Claude, paired with Akka Specify, handled implementation, testing, and review. This structured approach allowed the team to isolate specific failure points within the migration pipeline, providing a clearer view of where the automation succeeded and where it required human intervention. The process was not a simple find-and-replace operation but a complex negotiation between the original code logic and the new architectural constraints imposed by the target environment.
The results were mixed but significant.
Akka reported a lines of code or performance improvement in 57 of the 65 ports. However, the company noted that gaps in context files remained around cross-component decisions. This suggests that while AI can handle isolated components, it struggles with holistic system understanding without precise, typed behavioral definitions. The quality of the specification matters more than the raw power of the model, a point that resonates deeply with teams considering AI-assisted modernization. It implies that the effort invested in defining clear behavioral constraints will yield better first-pass implementations than relying solely on model capability. This finding is vital for engineering leaders who are trying to balance speed with reliability in their migration roadmaps.
"Structured specifications with claims, evidence, and typed behavior improved first-pass implementations," the report says.
Model efficiency and cost
Model choice impacted both speed and cost. According to the InfoQ report, Sonnet averaged 61 minutes per port versus 120 minutes for Opus. However, Opus used about 40% fewer tokens. Higher effort settings increased consumption without consistently improving efficiency. This trade-off highlights the economic reality of using frontier models for bulk code migration.
Commenters on LinkedIn noted that smaller models followed the spec more closely, while larger models tended to improvise. Aaditya, a commenter cited by InfoQ, wrote: "Smaller model's behavior matched modernization work he had observed, where the small model follows the spec while a larger model may improvise." This observation aligns with the token usage data, suggesting that constraints can yield better results than raw capability. The implication for enterprise teams is that choosing the largest model available is not always the optimal strategy for cost-effective migration projects.
For teams managing budgets, this data provides a baseline.
If you are porting legacy code, the spec quality may determine whether you save time or waste compute. The 9.41 billion token figure is a hard number to plan around. It is not a trivial cost, even for a large enterprise.
Open source governance in the AI era
This experiment occurs against a backdrop of shifting norms in open source security and maintenance. On October 1, Google suspended product vulnerability submissions to its Open Source Software Vulnerability Reward Program due to an influx of invalid AI-driven reports, according to Tom's Hardware. The program will not resume until the first quarter of 2027. This move highlights the strain that AI-generated content is placing on human-moderated systems.
Similarly, Anil Madhavapeddy, a professor at Cambridge, argued in an article discussed by InfoQ that AI agents are disrupting traditional security disclosure. He described how probes appeared in his webserver logs minutes after he opened a PR to fix a vulnerability. "It looks to me like our security processes need to invert somewhat," he wrote. The time between disclosure and exploitation is shrinking as AI agents can generate exploits from limited clues.
These governance issues are not abstract. They affect the same open source projects that AI agents are now being asked to port and maintain. If AI can write the code, it can also find the bugs. The challenge is distinguishing helpful automation from malicious or sloppy generation. The Akka experiment attempts to solve the former; the Google and Cambridge reports illustrate the latter.
Infrastructure and tooling
The rise of AI-assisted development is also changing the infrastructure surrounding open source. GitHub released ReviewBench, an open benchmark for AI code review, on October 5. According to the GitHub blog, the benchmark uses 219 public pull requests across 19 languages, modeled after over 100 million real pull requests. It aims to provide a rigorous, reproducible evaluation methodology for code review agents.
Meanwhile, Reflection AI released Beam, an open-weight model with 501 billion total parameters, on October 5. According to Semafor, Beam is positioned as a "workhorse" for coding and agentic tasks, requiring three to four times less computing power than comparable open models. The release reflects a trend toward efficient, open-weight models that can be deployed locally or in sovereign AI systems.
These tools form the new stack for open source development.
You have the models (Beam, Sonnet, Opus), the benchmarks (ReviewBench), the workflows (Akka Specify), and the security challenges (Google VRP suspension). The ecosystem is maturing, but it is doing so under pressure. The 9.41 billion tokens used by Akka are a reminder that this is not a free lunch. It is a resource-intensive process that requires careful management.
The takeaway for engineers is clear. AI can port code, but it needs structure. It can review code, but it needs benchmarks. It can find vulnerabilities, but it can also flood them. The direction of open source is not just about who writes the code, but how we verify, secure, and maintain it in an age of synthetic generation.
Sources
8- 01Akka Tests Spec-Driven AI Delivery across 65 Open Source ProjectsEN
- 02Google freezes open-source bug bounty programEN
- 03AI Agents Are Disrupting Open Source Security DisclosureEN
- 04ReviewBench: An open benchmark for AI code reviewEN
- 05Reflection AI unveils an open-source Western answer to Chinese labsEN
- 06Open Source as We Know It Is DeadEN
- 07Open source didn't get less important because AI can write the codeEN
- 08DigitalOcean Ends Open Source Credits ProgramEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.