Monty: Secure Python Sandbox in Rust for AI Code
What Monty is
Monty is a minimal, secure Python sandbox written in Rust, built by the Pydantic team for running code written by AI. The motivation is a pattern that has spread across agent frameworks: rather than making a long sequence of individual tool calls, a model writes a short program that calls your tools, does arithmetic, and prints a result. The Monty docs cite Cloudflare's code mode, Anthropic's programmatic tool calling and code execution with MCP, and Hugging Face's smolagents as examples. All of them need somewhere safe to run that generated code, and Monty is meant to be that place.
Its pitch is avoiding the latency, complexity, and cost of a container-based sandbox. Inside a Monty sandbox there is no filesystem, no environment variables, and no network. Code reaches the host only through the functions and mounts you explicitly pass in. Monty also powers Code Mode in Pydantic AI.
It ships in two forms. OSS Monty is the MIT-licensed sandbox you install as a package for Python, JavaScript/TypeScript, or Rust. Full Monty is a commercial server that runs the same sandbox behind a WebSocket as a container image, adding OS-level isolation and horizontal scaling. The docs announce that Monty has reached v1.0.0 and that the team now considers it ready for production use.
How it works
Monty is not a wrapper around CPython. It is its own interpreter for a deliberately limited subset of Python, implemented in Rust, with the goal that everything it implements behaves exactly like CPython 3.14 and every divergence is documented.
Applications create a pool of worker subprocesses up front. Getting a sandbox is a checkout from that pool, and each command sent to a session is one message each way. Sessions persist, so a REPL-style agent that runs ten snippets does not re-run earlier snippets for each new one.
Host interaction is explicit. In the README example, the model's code calls nutrition('chocolate bar'); that function is passed in through external_lookup, runs on the host, and the sandbox only sees its return value. Inputs such as bulb_watts are passed in as values.
Resource limits are enforced by the interpreter itself. The docs give the example that 'x' * 10**12 raises MemoryError before the allocation is attempted. Execution time is limited the same way.
Key features
- No ambient host access: filesystem, environment, and network do not exist inside the sandbox unless you pass in functions or mounts.
- Pooled, fast startup: the README claims a new OSS Monty sandbox takes under 1 ms from a running pool, and about 2 ms for Full Monty.
- Persistent sessions: REPL state survives between snippets, which suits multi-step agents.
- Snapshots: the whole sandbox state can be dumped to bytes at an external function call or at the end of a snippet and resumed later, which makes long external calls and human-in-the-loop pauses cheap.
- Strict resource limits: memory and execution time are enforced by the VM.
- Three host languages: packages on PyPI (
pydantic-monty), npm (@pydantic/monty), and crates.io, with community bindings for Go and Dart/Flutter. - Type checking: the docs recommend turning on type checking so unsupported APIs fail before they run.
- Async support:
async/awaitwithasyncio.run,asyncio.gatherfor concurrent host calls, andasyncio.sleep. - Documented divergences: one page per builtin, module, or construct listing every known difference from CPython 3.14.
Getting started
Install for your host language:
uv add pydantic-monty # Python
npm install @pydantic/monty # JavaScript / TypeScript
cargo add monty # RustThe README's Python example runs model-written code with one input and one host function:
from pydantic_monty import Monty
code = """
kcal = nutrition('chocolate bar')['kcal']
hours = kcal * 4184 / (bulb_watts * 3600)
print(f'a chocolate bar powers a {bulb_watts} W bulb for {hours:.1f} hours')
"""
with Monty() as pool:
with pool.checkout() as session:
session.feed_run(
code,
inputs={'bulb_watts': 10},
external_lookup={'nutrition': lambda food: {'kcal': 230}},
)
#> a chocolate bar powers a 10 W bulb for 26.7 hoursThe pool is created once and reused; each checkout() gives a fresh session. The docs include matching quickstarts for JavaScript and Rust.
Use cases
- Code mode for agents: let a model write a short script that calls several tools and combines their results, instead of round-tripping each tool call through the model.
- High-volume agent backends: because a sandbox is a pool checkout rather than a new container or VM, the docs argue you can run thousands of workers with minimal cost and complexity.
- Human-in-the-loop workflows: snapshot a sandbox when it calls an external function that needs approval, store the bytes, and resume when a person responds.
- Long-running REPL sessions: persist state between turns of a conversation without keeping a process alive.
- Safe calculation: give a model a place to do arithmetic and data shaping it should not do in its head.
How it compares
The project's docs compare Monty with Docker, Pyodide, WASI, and hosted sandboxing services. In its own measurements of creating a sandbox and running ten REPL commands, the docs report about 1.2 ms for OSS Monty, 7 ms for Full Monty over WebSocket, 200 ms for WASI with wasmtime, 900 ms for local Docker, 1,900 ms for a sandboxing service, and 2,700 ms for Pyodide in Deno. These are Pydantic's figures, and part of the gap comes from Monty pooling and keeping sessions while the others are measured starting from nothing and re-running earlier commands.
The trade-off is language coverage. Docker, sandboxing services, and Pyodide run real CPython with its standard library and third-party packages. Monty runs a subset with no third-party packages. The docs also point out that OSS Monty runs code on the same machine as your application, while Full Monty and hosted services run it remotely, which reduces the blast radius of an escape.
Things to know before adopting
- It is a Python subset: class inheritance, metaclasses, generators (
yield),match,del, method decorators such as@property, user-defined exceptions, and wildcard imports are not supported. Many are rejected at parse time withNotImplementedError. - Limited standard library: modules such as
json,re,datetime,math, anditertoolsexist, often partially;enum,logging,hashlib,uuid,urllib, and others are absent, andsubprocess,socket, andthreadingare excluded by design. - Behavioral differences:
map,zip,filter, and similar builtins are eager,reis backed by Rust's fancy-regex, and only UTF-8, ASCII, UTF-16, and UTF-32 codecs exist. - Open-core split: OSS Monty is MIT; Full Monty, with OS-level isolation and remote execution, is commercial.
- Security scope: the docs have a dedicated security model page explaining what "secure" does and does not mean; read it before running untrusted code in-process.
Project activity
As of October 2026 the repository has roughly 8,500 stars. The GitHub repository was created on May 28, 2023, the project is written in Rust, and it is licensed under MIT. Its documentation announces the v1.0.0 stable release. The source is at github.com/pydantic/monty, and documentation is at pydantic.dev/docs/monty. CI, coverage, and CodSpeed performance tracking badges appear in the README, and the project sits alongside Pydantic AI, Pydantic Logfire, and the Logfire AI Gateway in the Pydantic stack.
Enjoying this project?
Discover more amazing open-source projects on TechLogHub. We curate the best developer tools and projects.
Repository:https://github.com/pydantic/monty
GitHub - pydantic/monty: Monty: Secure Python Sandbox in Rust for AI Code
Monty is Pydantic's MIT-licensed Python sandbox written in Rust for running code generated by LLMs. It implements a Python 3.14 subset with no filesystem, envir...
github - pydantic/monty
Related Projects
Docling: Document Parsing for Generative AI
Python toolkit that converts PDFs, Office docs, HTML, audio and more into structured Markdown or JSON for AI pipelines.
Hyperswitch: Open-Source Payments Orchestration in Rust
Juspay's modular, Rust-based payments switch with multi-PSP routing, retries, vault and reconciliation.
PageIndex: Vectorless, Reasoning-Based RAG
Vectorless RAG: builds a tree index of each document and lets an LLM reason its way to the right section.

