Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

ServiceNow's AutoSynthData Turns Agent Failures Into Training Data

ServiceNow published an enterprise agent pipeline on 2 October that generates new training tasks from a model's own failures, targeting the retrieval and tool-use gaps that block deployments. AutoSynthData was detailed on Hugging Face by four ServiceNow researchers.

AI & modelsNewsRachel NwosuPublished: 2 October 20267 min readSources 14
ServiceNow's AutoSynthData Turns Agent Failures Into Training Data

The pipeline is called AutoSynthData. It uses a target model's failures and a stronger teacher's successes to decide what the model should learn next, then generates and validates new tasks that exercise those capabilities, according to the engineering post published on 2 October.

That matters because the enterprise agent problem is largely a retrieval problem. An agent that cannot find the right policy, asset record or work note at the right moment fails the task, and a single failure is not enough to train on.

The ServiceNow authors, Esakkivel Esakkiraja, Shruthan Radhakrishna, Denis Akhiyarov and Sagar Davasam, describe the gap plainly: a model may be broadly capable and still struggle with a particular environment, a workflow it handles poorly, a combination of tools it misuses, or a constraint it fails to respect. The post sets three conditions for a generated task, feasibility, realism and difficulty, and three for the verifier that grades the result, starting with consistency. The verification layer is the load-bearing part: without a reliable check, synthetic tasks teach nothing. ServiceNow illustrates the method on EnterpriseOps Gym, a benchmark released with a paper this year. It is one of several research efforts landing in the same window that attack the same problem from different angles: how to ground enterprise systems in context they can actually act on.

Retrieval quality becomes a measurable variable

On 30 September, two preprint papers submitted to arXiv took aim at the representation layer beneath enterprise retrieval. Merieme Askour and Ayoub Merimi propose a theory of context sufficiency in generative AI personalization, arguing that once context is cheap to supply, more is not better.

Their framework identifies four states, insufficiency, sufficiency, saturation and interference, and defines a Context-Sufficiency Frontier to locate the minimal relevant set. In a full-factorial experiment with a generative recommender at a large home-furnishing retailer, relevant context improved appropriateness while irrelevant context reduced it and destabilized retrieval.

The second paper, from Terry Dorsey and Kevin Huggins, introduces Enterprise Representation Simplification and a representation-neutral model called Enterprise Representation Complexity. The authors argue that enterprise information accumulates structures shaped by applications, projects and organizational boundaries, and that this representational complexity must be maintained and interpreted by both people and AI systems. Their key claim for retrieval practitioners: reductions in task-level complexity reduce the representational extent an AI system must identify, relate and interpret, and text-to-SQL research provides evidence that reduced schema and reasoning complexity can improve accuracy. The authors are careful to note that ERC is not a universal complexity, performance or cost metric.

Neither paper ships a product. Both point at the same diagnosis, though. Throwing more context at a retrieval system is not the same as giving it sufficient context, and the difference shows up as unstable answers in production.

Identity is the other half of the grounding problem

The Hacker News published a framework on 28 September for identity and access management of AI agents, and it names a failure mode that retrieval teams tend to discover late. IAM platforms express intended access, while applications and infrastructure reveal what the agent actually executed.

Between the two sits what the article calls identity dark matter: the agents, credentials, application-local accounts and authentication paths that central identity data never reports. The piece cites OWASP's Top 10 for Large Language Model Applications, which names excessive agency as LLM06, and lists five common failures including absent ownership, long-lived secrets and unbounded delegation.

The retrieval angle is direct. An agent granted broad permissions can pull data it was never scoped to see, and static role assignment cannot bound that behavior. Configuration findings describe possibility; telemetry describes what occurred.

Cloud vendors push the plumbing down the stack

Microsoft made Azure Container Apps Express generally available, InfoQ reported on 1 October, alongside the general availability of Azure Container Apps Sandboxes, the isolated compute layer Express runs on. Express takes a container image, a region and whatever configuration the app needs, then provisions compute, ingress and scaling itself. It runs on consumption CPU with per-second billing and scales to zero when idle.

Microsoft describes Express as developer-first and agent-first, and its documentation states it is built entirely on Sandboxes, which provision from prewarmed pools for subsecond startup, isolate each workload in its own hardware-isolated microVM boundary, and support suspend and resume with sub-second restore. Developers can use Sandboxes directly, and Microsoft positions that route for agent platforms and secure code-execution services. Reaction on Reddit suggested the primitive was in use before the announcement; one commenter, MuhBlockchain, claimed ACA Sandboxes underpin core Azure services including Foundry Hosted Agents, a detail Microsoft's own material does not state.

CoreWeave moved in the same direction a day earlier. It unveiled CoreWeave Forge on 30 September at its Fully Connected event in San Francisco, an integrated software platform for developing, running and improving AI models and agents, according to Data Center Knowledge. Forge is available in Free, Pro and Enterprise editions and includes CoreWeave Notebooks, Agent Lens for observability of production agents, and RL Rollouts.

Corey Sanders, CoreWeave's senior vice president of product management, said at a media briefing that the company had moved from a focus on being the best provider for AI-centric infrastructure to delivering AI services that let customers build applications on its platform. IDC analyst Dave McCarthy told Data Center Knowledge that CoreWeave needs a wider enterprise audience and a bigger software ecosystem to build a sustainable business.

Where the demand is coming from

China's generative AI user base surpassed 700 million at the end of June, the China Internet Network Information Centre reported, up 16 per cent from 602 million at the end of 2025, when penetration stood at 42.8 per cent. The South China Morning Post, which covered the release on 29 September, noted that 76 per cent of surveyed users said they used generative AI to seek answers, 48 per cent to process images or videos, and 38 per cent for text processing.

Those are consumer numbers, not enterprise ones. But they set the expectation baseline that enterprise retrieval systems are now measured against, and they explain why vendors across the stack are racing to make agent outputs look less like guesses.

The infrastructure constraints underneath are real. Optical frequency comb generators, which produce many wavelengths from a single laser instead of one laser per channel, are being pitched as a way to cut power, cost and failure points as AI data centers scale past 16 wavelengths per fiber, Data Center Knowledge reported on 24 September. Frank Smyth, founder and CTO of Pilot Photonics, said comb lasers potentially allow many more wavelengths packed much tighter without fear of interference. Marcello Girardi, CEO and co-founder of Solinide, put the practical limit bluntly: four wavelengths is a solved problem for laser arrays, the difficulty starts at eight, and by 16 the component count limits you. Steven Estrella of Quintessent added that with the current indium phosphide CW DFB laser supply shortage, reducing the number of required lasers is more important than ever.

Enterprise buyers are also rethinking where inference runs. VMware's Private Cloud Outlook 2026, cited by Data Center Knowledge on 21 September, indicates that 83 per cent of enterprises have completed or are planning to repatriate workloads from the public cloud, driven by security, cost, compliance and performance concerns. Modern colocation facilities support 35 kW cabinets with air cooling and 70 to 150 kW cabinets with optional liquid cooling, which is where inference on confidential corporate data tends to land.

None of this solves the retrieval problem by itself. But the direction of travel across the last week is consistent: the industry is treating context, representation and identity as engineering variables rather than as prompts to be tuned. That is a harder problem than wiring up a vector store, and it is the one that decides whether an agent can be trusted with the work.

Comments 0

Sources

14
  1. 01AutoSynthData: Generating Training Data for Enterprise AgentsEN
  2. 02When More Data Is Not Enough: The Context-Sufficiency Frontier in Generative AI PersonalizationEN
  3. 03Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AIEN
  4. 04IAM for AI agents: A Practical Enterprise FrameworkEN
  5. 05Container Apps Express Reaches GA on a Newly Generally Available Sandbox LayerEN
  6. 06CoreWeave Targets Enterprises with Forge PlatformEN
  7. 07China's generative AI user base crosses 700 million, covering over half the populationEN
  8. 08The Role of Optical Frequency Comb Generators in AI Data CentersEN
  9. 09Enterprises Adopt Colocation for AI and Hybrid Cloud InitiativesEN
  10. 10Redefining enterprise intelligence with autonomous AIEN
  11. 11ShamAN-Q: Shampoo Augmented NanoQuant for Sub-1-bit LLM WeightsEN
  12. 12ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent RelationsEN
  13. 13Conformal Adversarial Generative EnsembleEN
  14. 14This startup helps food carts switch loud, dirty generators for batteriesEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.