Agentic AI meets the identity gap: the enterprise RAG problem no one solved
China's generative AI user base passed 700 million at the end of June, up 16 per cent from 602 million at the end of 2025, according to data released by the China Internet Network Information Centre on 29 September. But for enterprises building retrieval augmented generation systems, the harder problem is not adoption. It is governance.

The China Internet Network Information Centre put the country's generative AI user base at more than 700 million at the end of June, a penetration rate above half the population. That figure, reported by the South China Morning Post on 29 September, is an increase of 16 per cent from the end of 2025, when China had 602 million generative AI users and a penetration rate of 42.8 per cent. The same survey found 76 per cent of users turn to AI for answers, 48 per cent for images or video, 38 per cent for text processing and 33 per cent for work summaries, meeting notes or presentations.
Those numbers describe consumer behaviour. They say nothing about whether the systems answering those queries can be trusted inside a company.
That trust problem is now the centre of the enterprise RAG conversation. Retrieval augmented generation, the technique of grounding a language model's output in documents fetched at query time, has become the default architecture for corporate AI. Building a working prototype takes days. Making it reliable enough to run a business process is a different engineering discipline, as VentureBeat noted on 27 September. The gap between the two is where most deployments stall.
Who owns the agent?
The Hacker News published a framework on 28 September that treats each AI agent as a non-human identity with a human owner, a defined purpose, scoped authorization, an expiration and continuous monitoring. The article names five failure modes that conventional identity and access management was never built to catch: absent ownership, long-lived secrets, unbounded delegation, invisible instantiation and no expiration. It cites the OWASP Top 10 for Large Language Model Applications, which lists excessive agency as LLM06, meaning an agent granted broad permissions or autonomy exercises capability beyond its approved task.
The distinction the piece draws is between configuration and telemetry. Static role assignment cannot bound what an agent does at runtime, and reviewing configuration cannot measure it. Configuration findings describe possibility; telemetry describes what actually occurred. That is a direct challenge to the way most enterprises currently audit their AI systems.
The scale of the governance problem is visible in the numbers. Foundry's 2026 Cloud Computing Study found 74 per cent of enterprises accelerated cloud migrations last year, according to Data Center Knowledge. VMware's Private Cloud Outlook 2026 indicates 83 per cent of enterprises have completed or are planning to repatriate workloads from the public cloud, driven by security, cost, compliance and performance concerns. Those workloads increasingly include AI inference, which Data Center Knowledge reported is moving into colocation facilities where operators can offer 35 kW cabinets with air cooling and 70 to 150 kW cabinets with liquid cooling.
The reason is not only density. Enterprises in regulated industries prefer to run inference on confidential financial, operational and research data inside private colocation suites with dedicated network circuits and self-operated firewalls rather than sending that data to third-party AI startups. That preference has a direct consequence for RAG architectures, because retrieval means the system must touch the corporate corpus on every query.
Infrastructure vendors move in
Infrastructure vendors are repositioning around that requirement. CoreWeave unveiled Forge on 30 September at its Fully Connected event in San Francisco, an integrated software platform for developing, running and improving AI models and agents. Corey Sanders, CoreWeave's senior vice president of product management, said at a media briefing that the company had moved from a focus on being the best provider for AI-centric infrastructure to delivering AI services so customers can build applications on its platform. Forge ships in Free, Pro and Enterprise editions and includes CoreWeave Notebooks, Agent Lens for observability over production agents, and RL Rollouts. Early partners include VAST Data, CrowdStrike and ClickHouse.
IDC analyst Dave McCarthy told Data Center Knowledge that CoreWeave is trying to separate itself from specialized AI cloud providers such as Nebius, Lambda and Vultr while competing with AWS, Google Cloud and Microsoft Azure. His argument is that winning AI labs is not enough for a sustainable business; enterprises need a bigger software ecosystem and capabilities that feel turnkey rather than do-it-yourself.
Microsoft took the opposite route on the same day. Azure Container Apps Express reached general availability, a deployment model that removes the environment provisioning step and replaces most configuration with opinionated defaults, according to InfoQ. It ships alongside the general availability of Azure Container Apps Sandboxes, the isolated compute layer Express runs on. Sandboxes provision from prewarmed pools for subsecond startup, isolate each workload in its own hardware-isolated microVM boundary, and support suspend and resume with full state snapshotting. Microsoft's documentation, quoted by InfoQ, states that Express is built entirely on that layer and describes the product as developer-first and agent-first.
AI-assisted workflows can create and update apps far faster than anyone can configure infrastructure by hand.
That sentence, from Microsoft's own documentation, is the clearest statement of why the identity problem has become urgent. If agents can create and modify infrastructure faster than humans can review it, then the governance layer has to be automated too. InfoQ noted that reaction on Reddit suggested the Sandboxes primitive was in use well before the announcement, with one commenter claiming it underpins core Azure services including Foundry Hosted Agents, a detail Microsoft's own material does not state.
Research pushes on both ends
The research literature published in the last week shows the same tension from different angles. A paper submitted to arXiv on 29 September by Yining Lu, Aurelie Lozano, Xi Yang, Naoki Abe, Yu Deng and Meng Jiang proposes MoFlow, a method for generating agentic workflows that jointly optimize accuracy, cost, latency, robustness and consistency rather than committing to one fixed trade-off. The authors formulate the problem as a multi-objective Markov decision process and solve it with Convex-Hull Monte Carlo Tree Search. Evaluated against six baselines on six benchmarks spanning mathematics, code and question answering, MoFlow achieved the highest average hypervolume, though the authors concede that comparing against single-scalar optimizers is difficult because those baselines were rerun for each testing preference while MoFlow never saw them.
Another arXiv paper, submitted on 30 September by Tianyu Chen, Mingyuan Zhou and Jiaxing Wu, attacks the retrieval step itself. Their VHOP-Router trains an embedding model into an autoregressive multi-step retriever that operates directly in visual latent space, retrieving linked image chains in a single tool call without the agent formulating intermediate text queries. The reported results are large: retrieval performance rose from under 5 per cent to 76.3 per cent, task success rates improved by 52.7 per cent, average token length fell 61 per cent from 1886 to 728, in-context images dropped by a factor of 23 and cumulative API payload by a factor of 35. Upgrading the agent itself yielded only a 3.7 per cent gain by comparison.
A third paper, submitted on 29 September by Sumit Asthana, Michael Ion and Kevyn Collins Thompson, addresses the synthetic data used to fine-tune conversational systems. The authors argue that prompting LLMs directly yields low-diversity data that collapses onto dominant modes, and propose training Generative Flow Networks on latent conversation structure using a Gaussian mixture density over interaction features. Across tutoring and emotional support dialogues, they report a better balance of fidelity, mode coverage and authenticity than reinforcement-learning and end-to-end LLM baselines, without copying training data.
Two more arXiv submissions from the same window are narrower but relevant to anyone running retrieval at scale. A paper submitted on 26 September by Junoh Kang, Kiseop Lee and Bohyung Han describes ReLOBGen, a method for generating limit order book messages that are replayable by construction, achieving 100 per cent replayability in 500-message rollouts and a 2.7 to 3.6 times speedup per replayed message over the LOBS5 baseline. A paper submitted on 28 September by Julien Moreau and Marc Lelarge proposes EnJoi, a diffusion-based data assimilation algorithm that learns the joint distribution of past and future states, with improved reconstruction on fluid and traffic flow simulations where observations are sparse and non-homogeneous.
None of these papers mentions enterprise RAG directly. Their relevance is structural: each one tries to make a generative system's behaviour more predictable, more measurable or more controllable, which is exactly what the identity framework published by The Hacker News says is missing from most corporate deployments.
The numbers underneath
Hardware pressure is adding to the problem. Data Center Knowledge reported on 24 September that optical frequency comb generators, which produce many evenly spaced wavelengths from a single pump laser rather than an array of individual lasers, are being evaluated for AI data center links. Marcello Girardi, chief executive and co-founder of Solinide, told the publication that four wavelengths is a solved problem for laser arrays, the difficulty starts at eight, and by 16 the component count limits you. Steven Estrella, director of product management at Quintessent, said combs reduce the number of lasers required for the same channel count, which translates to fewer points of failure and less wavelength control complexity, and noted that with the current indium phosphide continuous-wave DFB laser supply shortage, reducing the number of required lasers matters more than ever.
The environmental ledger is also being counted. A study by the Fraunhofer Center for Silicon Photovoltaics, commissioned by the German Mineral Resources Agency and reported by pv magazine on 1 October, found that modules installed in Germany by the end of 2023 represented a material stock of around 5.35 million metric tons, including 3.85 million tons of glass, 663,000 tons of aluminium, 208,000 tons of silicon, 53,500 tons of copper and 3,900 tons of silver. The potential material value is €4.8 billion. More than 600,000 tonnes of PV waste could arise by 2030, but the recorded volume of end-of-life modules declined between 2022 and 2024, contrary to expectations, with researchers citing export of used modules as one possible reason. In 2024, more than 50,000 tonnes of recycling capacity was available while only around 9,300 tonnes underwent initial processing.
That mismatch between installed capacity and actual throughput has an obvious parallel in enterprise AI. Capability is being deployed faster than the systems to govern, measure and retire it.
Where the two stories meet
Generative Bionics, the startup spun out of the Istituto Italiano di Tecnologia, offers a concrete example of how quickly commitments accumulate. Il Sole 24 ORE reported on 1 October that the company closed a €70 million round on 4 December 2025, led by Cdp Venture Capital with participation from Eni Next among others, became operational on 7 January and presented its robot Gene.01 in July. It has a four-year programme with Fincantieri for a welding humanoid, with first employment expected in early 2027, and a newly signed agreement with Eni to evaluate industrial collaborations starting from Gene.01, covering inspection, teleoperation, remote assistance and support in complex environments.
Chief executive Daniele Pucci told Il Sole 24 ORE that a robot that does everything does everything badly, describing the company's approach as a common platform customized at three levels: the AI model, the external design, and the hands and feet. The company says it launched that strategy first in February and now sees others following it.
Every one of those deployment scenarios, from welding to hospital logistics, generates the same question the identity framework raises: who is accountable when the agent acts outside its approved task? The answer, according to the framework, cannot come from configuration review alone. It has to come from runtime telemetry, and that is a capability most enterprises have not yet built.
As demand elsewhere shows no sign of slowing. PopWheels, a Brooklyn startup, is testing a battery-swapping model for New York food carts that replaces fossil-fuel generators, running a six-month pilot launched in July with two charging cabinets in Flushing Meadows Corona Park and 10 participating vendors, backed by $46,000 from the City Parks Foundation and $30,900 from Resilient Cities Catalyst, according to Canary Media. Vendor William Arevalo told the publication his generator costs roughly $500 to $600 a month to fuel. The company already charges delivery workers between $65 and $95 a month for e-bike battery access from more than 45 cabinets across the city.
Cheaper, cleaner infrastructure keeps arriving. The governance layer keeps lagging.
Sources
14- 01China's generative AI user base crosses 700 million, covering over half the populationEN
- 02IAM for AI agents: A Practical Enterprise FrameworkEN
- 03CoreWeave Targets Enterprises with Forge PlatformEN
- 04Container Apps Express Reaches GA on a Newly Generally Available Sandbox LayerEN
- 05Enterprises Adopt Colocation for AI and Hybrid Cloud InitiativesEN
- 06The Role of Optical Frequency Comb Generators in AI Data CentersEN
- 07MoFlow: Multi-Objective Agentic Workflow GenerationEN
- 08Learning to Route in Visual Space via Multi-Step Embedding RetrievalEN
- 09Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow NetworksEN
- 10ReLOBGen: Replayable Limit Order Book Message GenerationEN
- 11EnJoi: Ensemble Joint Score Filter for Generative Data AssimilationEN
- 12Germany could generate over 600,000 tonnes of PV module waste by 2030EN
- 13This startup helps food carts switch loud, dirty generators for batteriesEN
- 14Generative Bionics, la start up che in meno di un anno ha stretto accordi con Eni e FincantieriEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.