OpenShell: Policy-Enforced Sandbox Runtime for AI Agents
GitHub Repo
Apache-2.0
October 2, 2026 at 09:19 AM
0 views

OpenShell: Policy-Enforced Sandbox Runtime for AI Agents

@NVIDIAProject Author

What OpenShell is

OpenShell is an open-source runtime from NVIDIA for running autonomous AI agents with bounded access. The README describes it as a safe, private runtime for fleets of autonomous agents. The premise is practical: agents are most useful when they can read files, install packages, call APIs, and use credentials, but handing them a normal shell gives them unrestricted access to your data, secrets, and network. OpenShell lets agents do that work inside a sandbox, and you declare in a policy what each agent is allowed to touch. Anything the policy does not allow is denied.

It is aimed at developers running coding agents locally, platform teams hosting agents for many users, and security teams that need a reviewable record of what an agent can reach. It is written in Rust, licensed under Apache-2.0, and at the time of writing is on its 0.1.x release line, which the README says introduced a stable release cadence, new isolation primitives, an expanded extension surface, and new APIs.

How it works

OpenShell governs agents in two ways. First, it instruments the kernel to enforce policy on every file access, system call, and network connection at runtime. Second, it uses formal verification to check what a proposed policy change would allow before that change is applied.

Gateway, supervisor, and sandbox

The gateway is the control plane. It authenticates users, stores sandbox state, delivers policy, attaches credential providers, and coordinates connections. When you create a sandbox, a compute driver provisions two things: the workload, where the agent runs, and a separate supervisor on the trusted side of the boundary. The supervisor confirms the isolation boundary is in place before the agent starts, then connects back to the gateway for policy, credentials, logs, and interactive sessions.

Inside the workload, a small sandbox component launches the agent as a child process. According to the architecture docs, on the current Linux backend the workload runs as a single non-root identity with no Linux capabilities, Landlock limits filesystem access, and seccomp user notification intercepts network operations. The sandbox never makes policy decisions; it reports what the agent is trying to do, identifies the calling executable from trusted /proc data, and hands the request to the supervisor.

The network path

When the agent opens a TCP connection or performs a DNS lookup, the sandbox forwards it over a mutually authenticated HTTP/2 channel to the supervisor, which checks it against policy, adds any credentials the policy allows, and only then opens the real connection. An outer network fence denies every other egress path. Each runtime builds that fence with its own tools: Docker and Podman turn off networking in the workload container and use a Unix socket, Kubernetes uses a NetworkPolicy plus mutual TLS, and the VM runtime gives the guest no network device and talks over vsock. If the supervisor disconnects, the agent is frozen until the same supervisor reconnects.

Credentials

Agents never see real credentials. A provider maps a service name to a stored credential, and the supervisor attaches it only to requests bound for endpoints the policy approves.

Key features

  • Declarative YAML policies: sections for filesystem_policy, landlock, process, network_policies, and network_middlewares.
  • Per-binary network rules: rules list which destinations each binary may reach and can restrict requests, for example allowing reads from an API but not writes.
  • Live network policy updates: filesystem and process settings are fixed at startup, but network rules and middleware can change while the sandbox runs.
  • Policy prover: formal verification flags risky changes such as new credentialed reach, new HTTP methods, or access to cloud metadata endpoints, and any finding blocks auto-approval. It also ships as a standalone openshell-prover command for CI.
  • Policy advisor: an agent can propose narrow network rules for human review as it discovers what it needs.
  • Global policies: a gateway administrator can apply one policy that replaces every sandbox's own policy.
  • Multiple runtimes: Docker, Podman, Kubernetes via Helm, and MicroVMs, behind one isolation backend interface.
  • Per-sandbox credentials: the gateway issues JWTs scoped to one sandbox and one run, rotated on restart.
  • SDKs: Python, TypeScript, Go, and Rust clients for the gateway.
  • Agent skills: installable skills that teach coding agents to drive the CLI, write policies, and debug gateways.
  • Extensibility: middleware, interceptors, and compute drivers.

Getting started

The README lists Linux, macOS on Apple Silicon, or Windows with WSL 2 (experimental), plus Docker, Podman, or host virtualization. Install the CLI and a local gateway, then create a sandbox:

curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
openshell sandbox create --name demo

The default sandbox image is a minimal Ubuntu with no agent installed. The "Run Your First Agent" guide in the docs runs OpenCode against a free OpenRouter model and walks through approving new access as the agent asks for it. To give your own coding agent OpenShell skills:

npx skills add NVIDIA/OpenShell

For application code that talks to a gateway, the Python SDK installs with:

uv add openshell

Use cases

  • Running coding agents on a laptop: let an agent install packages and run builds in a sandbox that can only reach the package registry and the model API you allow.
  • Keeping API keys away from agents: store a GitHub or inference credential as a provider so the agent can call the API without the token ever entering the sandbox.
  • Shared agent platforms on Kubernetes: deploy the gateway with Helm and give each user isolated sandboxes, with a global policy as an organization-wide baseline.
  • Reviewing access growth: route agent-proposed network rules through the prover and advisor so only flagged changes need a human.
  • Policy checks in CI: run openshell-prover to confirm a policy grants no more access than a defined boundary before it ships.
  • GPU workloads: the sandbox docs cover images, runtimes, and GPUs.

How it compares

The README does not compare OpenShell to named alternatives. Conceptually it differs from running an agent in a plain container: a container with networking enabled can usually reach anything, while OpenShell's design makes the supervisor the only egress path and evaluates every connection against policy per binary. It also differs from sandboxes that only isolate code execution, since OpenShell additionally brokers credentials and verifies policy changes. The trade-off is more moving parts, including a gateway, a supervisor per sandbox, and runtime drivers, to install and operate.

Things to know before adopting

  • Early version numbers: the project is on 0.1.x and the repository was created in February 2026. Read the 0.1.0 upgrade guide and expect API evolution.
  • Platform support: Linux, Apple Silicon macOS, and experimental Windows via WSL 2. On Kubernetes, your CNI must enforce NetworkPolicy, or the network fence does not hold.
  • Telemetry is on by default: the README says it collects anonymous operational categories and counts, not sandbox names, hostnames, paths, prompts, credentials, model names, or user content. Disable it with OPENSHELL_TELEMETRY_ENABLED=false on the gateway, server.telemetryEnabled=false for Helm, or compile it out.
  • Shared deployments need token expiry: the docs recommend setting gateway_jwt.ttl_secs on Kubernetes; local single-user gateways can leave it unset.
  • External materials: the README's disclaimer notes the software retrieves external materials governed by their own licenses.
  • License: Apache-2.0.

Project activity

As of October 2026 the repository has roughly 14,200 stars. It was created on February 24, 2026, is written in Rust, and is licensed under Apache-2.0. The source is at github.com/NVIDIA/OpenShell, and documentation is at docs.nvidia.com/openshell/latest. The project publishes a public roadmap and RFC board, takes questions through GitHub Discussions, publishes community telemetry reports, and describes itself as built agent-first, developed with the same agent-driven workflows it enables.

Enjoying this project?

Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.

Project
nvidia-openshell
Created
October 2
Last Updated
October 2, 2026 at 09:19 AM

Find more projects like this

One email a week: new and trending developer tools, fresh comparisons, and what shipped. Unsubscribe in one click.