Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

AI agent tooling grows up: runtimes, registries and scanners arrive

A cluster of open-source projects landing on Hacker News shows what enterprise AI agent tooling looks like once teams stop prototyping: runtimes, inboxes, artifact registries and security scanners. The most recent, Hyperlane, pitches an IDE that runs agents in parallel.

AI & modelsExplainerRachel NwosuPublished: 28 September 20265 min readSources 5
AI agent tooling grows up: runtimes, registries and scanners arrive

The pitch is no longer "build an agent". It is "keep a fleet of them alive, auditable and fed". Four projects posted to Hacker News over nine months sketch that shift: a runtime, an inbox, a registry and a scanner. None of them is a model. All of them assume you already have one.

Soma, documented at docs.trysoma.ai, is the most explicit about the plumbing. It is an open-source, self-hostable AI agent and workflow runtime that ships as a single binary and describes itself as a security and governance plane across your agents. Agents and MCP functions are written in ordinary code, TypeScript today with Python listed as coming soon, and the runtime handles fault tolerance and resumability: execution can crash or be suspended at any point and resume where it left off. It also auto-generates A2A endpoints, with OpenAI Streaming compatibility marked coming soon, and bundles an MCP server pre-integrated with third-party SaaS providers that manages credential encryption and rotation.

The docs list an outbound AI gateway that intercepts agent requests to model providers for observability, fine-grained API key access management, and local, AWS or soon GCP KMS encryption for secrets. Platform support is uneven. TypeScript and Python builds cover macOS x86 and ARM plus Linux GNU x86 and ARM, with Windows unsupported. Rust is listed as not yet available on any platform, which the docs attribute to the project's use of Unix domain sockets.

Agents that keep working after you close the tab

Pizza Bot attacks the same problem from the interface side. Its GitHub repository describes a local-first inbox for long-running AI agents, built with DeepAgents and LangGraph, where you start or schedule a task, walk away, and find finished work in Unread and approval decisions in Action. The README states agents keep working when you navigate away or disconnect, provided the api-server process stays running.

The architecture is a stateful DeepAgents/LangGraph runtime serving the same React experience in Electron and the browser, with desktop app, web app and terminal CLI all talking to the api-server over HTTP and SSE. Checkpointed runs survive client disconnects, and cron or webhook triggers can start work without an open conversation. The repository says Pizza Bot was developed at Amazon and is released under the Apache 2.0 license. It supports Amazon Bedrock, Anthropic, Google Gemini, OpenAI, OpenRouter and Ollama, and grants no default home-directory access: folders must be added individually as read-only or writable under Settings, Files.

Local-first by default. The api-server binds to 127.0.0.1; non-loopback binding requires authentication.

That line from the README matters more than the feature list. Agents that run unattended and hold credentials are a different security object from a chat window, and the projects that survive procurement will be the ones that say so on the first page.

The registry layer gets an open-source option

Artifact Keeper is not about agents at all, which is exactly why it belongs in this picture. It is a self-hosted artifact registry positioned as a drop-in replacement for JFrog Artifactory and Sonatype Nexus, with no feature gates and no open-core split. The backend is Rust with Axum, SQLx, PostgreSQL and Wasmtime, and the project claims native protocol support for more than 45 package formats rather than treating formats as labels on a blob store.

The repository describes three repository types matching the Artifactory and Nexus model: local, proxy and virtual. Proxy repositories cache artifacts from public registries such as npmjs.com, PyPI, Maven Central and Docker Hub on first request, while virtual repositories aggregate several repos behind one URL and resolve local packages first. Security tooling includes Trivy vulnerability detection, OWASP Dependency-Track for SBOM analysis, OpenSCAP compliance auditing and a policy engine with quarantine workflows and scan-before-download enforcement. Container images are built on DISA STIG-approved Red Hat UBI 9 base images with non-root execution, per the README.

Elsewhere the project lists GPG and PGP artifact signing with key management, a WASM plugin system for custom format handlers, mesh-based peer replication with label-based sync policies, and OpenID Connect, LDAP, SAML 2.0 and JWT authentication with per-repository RBAC. There is a built-in migration tool for moving repositories, artifacts, users and permissions off Artifactory, plus OpenSearch-powered full-text search and an OpenAPI 3.1 spec covering 277 operations. Installation is a git clone followed by docker compose up.

Someone has to audit the MCP servers

Golf Scanner exists because the first three projects all expand the attack surface. It is a free, open-source CLI, a single static binary written in pure Go with three dependencies, zero telemetry and no account required. It discovers MCP server configurations across seven IDEs: Claude Code, Cursor, VS Code, Windsurf, Gemini CLI, Kiro and Antigravity.

It then runs 20 security checks, nine offline and 11 online, the latter querying OSV, GitHub, npm, PyPI, OCI registries and the MCP Registry. Checks cover command safety, plaintext credentials in arguments, URLs and environment variables, script location and permissions, binary permissions, container isolation flags such as privileged mode, dangerous volume mounts including the Docker socket, typosquatting, digest pinning and cosign signature verification. Each server gets a 0 to 100 risk score, severity-weighted with critical findings counted at 10x and hard caps: any critical finding caps the score at 30, any high finding at 59.

The scanner ships a --fail-on flag that returns exit code 1 when findings meet a severity threshold, which is how it gets into CI. Most scans need no token, since the tool makes roughly three GitHub API calls per unique repository with results cached.

Read together, the four projects describe a stack rather than a trend: a runtime to execute agents, an inbox to supervise them, a registry to feed them packages, and a scanner to check what they were handed. All four are open source. The newest, Hyperlane, promises an IDE and ADE that merge agent worktrees with native tooling. Whether enterprises assemble this stack themselves or buy it assembled is the open question, and the governance headlines of late September suggest plenty of vendors are betting on the second option.

Comments 0

Sources

5
  1. 01Show HN: Hyperlane – A IDE and ADE merging agent worktrees with native toolingEN
  2. 02Show HN: Pizza Bot – An inbox for AI agents that work in the backgroundEN
  3. 03Show HN: I built an open-source Rust/TS AI agent runtime with a Next.js-style DXEN
  4. 04Show HN: Artifact Keeper – Open-Source Artifactory/Nexus Alternative in RustEN
  5. 05Show HN: Golf Scanner – OSS tool to find and audit every MCP serverEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.