Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Agent tooling splits in two: runtimes that execute, scanners that audit

Five open-source agent projects posted to Show HN show the same split: runtimes that execute agent work, and security tooling that watches it.

AI & modelsExplainerRachel NwosuPublished: 28 September 20264 min readSources 5
Agent tooling splits in two: runtimes that execute, scanners that audit

Agent tooling is settling into two jobs: running agents long enough to be useful, and finding out what they can reach on your machine. The five projects below, posted to Hacker News' Show HN, sit at different points on that line.

Runtimes

Pizza Bot is the clearest example of the first job. Its GitHub README describes a local-first inbox for long-running AI agents, built on DeepAgents and LangGraph. Start or schedule a task, walk away, return to finished output in an Unread queue, and find approval requests in a separate Action queue. Agents keep working when you navigate away or disconnect, and only the api-server process has to stay running.

The project was developed at Amazon and is released under the Apache 2.0 license, according to its repository. It supports Amazon Bedrock, Anthropic, Google Gemini, OpenAI, OpenRouter and Ollama. Two details matter for anyone evaluating it. First, checkpointed runs survive client disconnects, and cron or webhook triggers can start work without an open conversation. Second, local file access is opt-in: folders are added individually under Settings, and the README states Pizza Bot receives no default home-directory access.

Soma, documented at docs.trysoma.ai, takes a different angle. It is a self-hostable agent and workflow runtime that ships as a single binary and is aimed at startups through to enterprise, according to its documentation. Code is written in TypeScript, with Python listed as coming soon. The runtime handles the plumbing around it: fault-tolerant, resumable execution that can crash or suspend and pick up where it left off, plus automatically generated A2A and OpenAI Streaming compatible endpoints.

Its platform table is unusually candid. TypeScript and Python builds are marked available on macOS Intel and ARM and Linux x86 and ARM, with Windows listed as planned but not natively supported because of the project's use of Unix domain sockets in Rust. Rust itself is marked unavailable across every platform in the table.

Registries and the developer loop

Artifact Keeper is not an agent runtime at all, but it lands in the same conversation because agents need somewhere to pull packages from. The GitHub organisation page describes a self-hosted artifact registry positioned as a drop-in replacement for JFrog Artifactory and Sonatype Nexus, with no open-core split and no enterprise edition. It claims native protocol support for more than 45 package formats, so pip, npm, docker, cargo, helm and go talk to it directly rather than through a generic blob store.

The project ships a backend in Rust with Axum, SQLx and Wasmtime, a Next.js dashboard, native iOS and Android apps, and a Helm chart with Terraform modules. Security features include Trivy scanning, SBOM analysis through OWASP Dependency-Track, GPG and PGP signing, and SSO via OpenID Connect, LDAP, SAML 2.0 and JWT. A built-in migration tool moves repositories, artifacts, users and permissions from JFrog Artifactory. The quickstart is three commands: clone, cd, docker compose up.

Claumon is narrower but more pointed. Its README describes a local dashboard for Claude Code that forecasts usage limits from a published empirical-Bayes model, refit daily on the user's own history, with an 80 percent credible interval and an ETA to threshold. Rate-limit gauges come from the Claude OAuth usage API, which the README contrasts with trackers that estimate limits from local logs.

The project's argument for existing is that Anthropic's usage analytics dashboard is aimed at Team and Enterprise org admins, not individual Pro or Max subscribers, and that personal /usage pages show current standing with no history. Claumon also manages the ~/.claude directory: per-session costs, SQLite-backed daily aggregates over a 7 to 90 day window, running-process control, and a memory-file browser. It binds to IPv4 loopback only, and the README recommends an authenticated SSH tunnel rather than exposing port 3131.

The audit layer

Golf Scanner covers the second job. It is a single static Go binary, with three dependencies, that discovers MCP server configurations across seven IDEs: Claude Code, Cursor, VS Code, Windsurf, Gemini CLI, Kiro and Antigravity. It runs 20 security checks, nine offline and 11 online, querying OSV, GitHub, npm, PyPI, OCI registries and the MCP Registry, then produces a 0 to 100 risk score per server.

The scoring rules are published. The worst severity per check sets a 0 to 10 score, scores combine as a severity-weighted average (critical 10x, high 7.5x, medium 5x, note 1x), and hard caps apply: any critical finding caps the total at 30, any high finding at 59. The README claims zero telemetry and no account requirement. Online checks cover command safety, plaintext credentials, world-writable scripts and binaries, container isolation, typosquatting, digest pinning and cosign signature verification.

Comments 0

Sources

5
  1. 01Show HN: Pizza Bot – An inbox for AI agents that work in the backgroundEN
  2. 02Show HN: I built an open-source Rust/TS AI agent runtime with a Next.js-style DXEN
  3. 03Show HN: Artifact Keeper – Open-Source Artifactory/Nexus Alternative in RustEN
  4. 04Show HN: Golf Scanner – OSS tool to find and audit every MCP serverEN
  5. 05Show HN: Claumon – forecasting Claude Code usage limits with a Gamma processEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.