Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Enterprise AI agent tooling splits between orchestration and security

Enterprise AI agent platforms are consolidating around incumbent stacks, while security researchers warn that agent identity and code sandboxing remain largely unsolved, according to vendor and analyst material published through 2026.

AI & modelsAnalysisRachel NwosuPublished: 27 September 20264 min readSources 4
Enterprise AI agent tooling splits between orchestration and security

Two things are true about enterprise AI agent tooling in 2026. The platform market is maturing fast. The security layer underneath it is not. Most buying decisions now sit in the gap between those two facts.

Start with the platform side. A Bardeen review of enterprise AI agent platforms, published on 10 August 2026, frames the choice as a question of which stack a company already runs: Salesforce Agentforce 360 for Salesforce operations, Microsoft Copilot Studio for Microsoft 365 shops, Google Gemini Enterprise for Google Cloud, ServiceNow for the IT and HR workflows it already owns. The review names WIQ as the AI-native successor to task mining. WIQ watches how work happens and ranks what is worth automating. The review also describes a common platform shape: a builder, an orchestration layer, and governance covering identity, observability, and controls on what agents may touch.

Gartner tracks the category as AI agent development platforms.

Pricing is where the marketing gets thin. Salesforce publishes Flex Credits at $500 per 100,000, with a standard action burning 20 credits, roughly $0.10, plus $2 per customer-facing conversation and user add-ons at $125 per user per month, according to the Bardeen review. A Manus comparison, published on 24 July 2026, lists the same $0.10 per action and $2 per conversation figures. It notes that total cost of ownership may exceed the advertised entry price, because most serious deployments also need Data Cloud, implementation services, and clean data work. Microsoft Copilot Studio is included with M365 Copilot at $30 per user per month, the Manus guide says. ServiceNow is listed as custom quote.

Adoption claims need reading carefully. The Bardeen review says 90% of the Fortune 500 already use Copilot Studio, and that Salesforce says 12,000 customers run Agentforce 360. Those are vendor-reported numbers. The same review cites Salesforce claiming Reddit deflected 46% of support cases and OpenTable resolved 70% of inquiries autonomously, again vendor-reported.

What the security research says

The n8n report on enterprise AI agent development tools is blunt about the layer buyers tend to skip. Agent authentication and identity is almost universally absent across the market, it says. Most marketing also conflates an agent using an API key to call a model with an agent presenting its own credentials to a third-party service. Only Google, Langflow, Workato, CrewAI, Sim.ai, and Gumloop scored 2 on that measure in the report. Lineage, the ability to trace an agent back to a human identity, is essentially non-existent, with only Google, Workato, and Gumloop scoring anything. Secrets management is similarly thin: only Google, Sim.ai, and Gumloop scored 2, with Make and Retool scoring 1.

Sandboxing for untrusted, LLM-generated code is rare. Only a handful of vendors offer a sandbox as a security boundary, the n8n report says, and most of those rely on third-party services, most commonly E2B. CrewAI deprecated its native code execution service and suggested customers use E2B instead. The report also notes that no platform does both agent code execution and human-written code execution well, not even Google.

Writing a prompt asking the LLM not to hallucinate or disclose sensitive data does not qualify as a security feature.

That line, from the n8n report, is the sharpest summary of the problem. Evaluations are sometimes used as guardrails, for instance an LLM-as-judge checking whether an answer contains PII. The report distinguishes that from a deterministic regex rule that detects a social security number and replaces it.

The adoption picture is messier than the platform pitches

Manus cites IDC research published in 2025 finding that 88% of enterprise AI proofs of concept do not reach production, with integration, governance, and cost planning usually the deciding factors rather than the AI itself. The Liqteq guide cites Gartner's June 2025 warning that over 40% of agentic AI projects will be cancelled by the end of 2027 due to escalating costs, unclear ROI, and inadequate governance. It describes agent washing, vendors rebranding chatbots and RPA tools as agents, as rampant. It says only approximately 130 of thousands of vendors are building real agentic systems.

Put those two halves together and the buying question changes. The platform shortlist is now fairly predictable, and it maps to whichever stack a company already runs. The harder question is whether the chosen tool can prove which agent acted, under whose authority, and inside what boundary. On the evidence published so far, that is still the part most vendors have not built.

Comments 0

Sources

4
  1. 0110 Best Enterprise AI Agent Platforms (2026)EN
  2. 0210 Best AI Agents for Enterprise in 2026, Reviewed and ComparedEN
  3. 03Enterprise AI agent development tools 2026EN
  4. 04Enterprise AI Agents: Real Use Cases, Costs & ROI Guide 2026EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.