Sandbox providers

E2B is the fastest general-purpose microVM sandbox for agents, but Fly.io Machines wins on price and control if you can skip an agent SDK.

Providers
20
Fields compared
26
Source confidence
32%
Last verified
2026-07-23 (2mo ago)
Re-verified
every 30 days
Data confidence0% of 406 figures verified in the last 30d
141sourced of 406141stale, oldest 8mo ago
20 tools · verified 2mo ago
Fly.io Machines logoFly.io MachinesRaw Firecracker microVMs with a REST API, durable volumes and 35+ regions.80.8

Runtime

Cold start (ms)
900ms
Cold start (claimed) (ms)
300ms
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
Yes
Runtimes
any-oci-image
Preview URLs
Partial
Browser inside
Partial
File up/download
Partial
Sydney region
Yes
Concurrent limit (sandboxes)
500sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
Partial
SDKs
rest-apicli
Streaming output
Partial

Pricing

Price / hour (/hr)
$0.031/hr
Meter
per-second
Idle cost
free-when-stopped

Traction

GitHub stars
·
Funding
series-c-plus

Positioning

Positioning
Hardware-virtualised containers that run anywhere, started and stopped over an API.

Timeline

Shipped
2022
Blaxel logoBlaxelAgent-first cloud claiming ~25 ms microVM boots from snapshots.79.6

Runtime

Cold start (ms)
500ms
Cold start (claimed) (ms)
25ms
Max runtime (ms)
1440min
Persistent FS
Partial
Snapshot & fork
Yes
Custom images
Yes
GPU
Unknown
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
Partial
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
Yes
SDKs
pythontypescript+1
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
·
Funding
seed

Positioning

Positioning
An agent-native cloud: sandboxes, agent hosting and a model gateway in one platform.

Timeline

Shipped
2025
Daytona logoDaytonaSub-second container sandboxes for agent workloads, from a team that built a dev-env manager.79.4

Runtime

Cold start (ms)
350ms
Cold start (claimed) (ms)
90ms
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
Unknown
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
container
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
Yes
SDKs
pythontypescript+1
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.17/hr
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
Funding
seed

Positioning

Positioning
Secure, elastic infrastructure for running AI-generated code.

Timeline

Shipped
2025
Runloop logoRunloopDevboxes, blueprints and eval harnesses aimed at teams training coding agents.79.0

Runtime

Cold start (ms)
800ms
Cold start (claimed) (ms)
1s
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
Partial
SDKs
pythontypescript+2
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
·
Funding
seed

Positioning

Positioning
Infrastructure and evaluation harnesses for teams building coding agents.

Timeline

Shipped
2024
Modal logoModalServerless Python compute with a Sandbox API bolted onto a GPU-first platform.77.5

Runtime

Cold start (ms)
1.5s
Cold start (claimed) (ms)
1s
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
Yes
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
gvisor
Root in guest
Yes
Egress control
proxy-logged

Agent integration

MCP server
Partial
SDKs
pythoncli
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.238/hr
Meter
per-second
Idle cost
active-cpu-only

Traction

GitHub stars
·
Funding
series-b

Positioning

Positioning
Serverless compute for AI teams, with sandboxes as one primitive.

Timeline

Shipped
2024
Vercel Sandbox logoVercel SandboxFirecracker microVMs for running untrusted code from inside a Vercel deployment.74.6

Runtime

Cold start (ms)
1.8s
Cold start (claimed) (ms)
Max runtime (ms)
1440min
Persistent FS
Partial
Snapshot & fork
Partial
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
No
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
2,000sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
No
SDKs
typescript
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.341/hr
Meter
per-second
Idle cost
active-cpu-only

Traction

GitHub stars
·
Funding
series-c-plus

Positioning

Positioning
Run untrusted code from your Vercel app in an ephemeral microVM.

Timeline

Shipped
2025
CodeSandbox SDK logoCodeSandbox SDKFirecracker VMs with memory snapshots, from the online IDE, now owned by Together AI.74.4

Runtime

Cold start (ms)
1.2s
Cold start (claimed) (ms)
1s
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
No
SDKs
typescriptpython
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
13,000
Funding
acquired

Positioning

Positioning
The VM infrastructure behind CodeSandbox, sold as an SDK.

Timeline

Shipped
2024
E2B logoE2BOpen-source Firecracker sandboxes with Python and TypeScript SDKs for AI agents.74.4

Runtime

Cold start (ms)
400ms
Cold start (claimed) (ms)
200ms
Max runtime (ms)
1440min
Persistent FS
Partial
Snapshot & fork
Partial
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescript+2
Preview URLs
Yes
Browser inside
Yes
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
Partial
SDKs
pythontypescript+1
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.17/hr
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
9,000
Funding
series-a

Positioning

Positioning
Open-source infrastructure for running AI-generated code.

Timeline

Shipped
2023
Namespace logoNamespaceFast microVM instances and CI runners with instant snapshots and cached builds.72.6

Runtime

Cold start (ms)
1.5s
Cold start (claimed) (ms)
800ms
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
Unknown
Runtimes
any-oci-image
Preview URLs
Partial
Browser inside
Partial
File up/download
Partial
Sydney region
No
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
proxy-logged

Agent integration

MCP server
No
SDKs
gorest-api+1
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-minute
Idle cost
free-when-stopped

Traction

GitHub stars
·
Funding
seed

Positioning

Positioning
Fast ephemeral compute for CI, builds and remote development.

Timeline

Shipped
2023
Islo logoIsloPer-agent isolated cloud sandboxes with enterprise policy controls, from Incredibuild.65.8

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Unknown
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Partial
Browser inside
Unknown
File up/download
Partial
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
Idle cost

Traction

GitHub stars
·
Funding
·

Positioning

Positioning
An isolated cloud environment per coding agent, with enterprise policy control over what it can reach.

Timeline

Shipped
·
Freestyle logoFreestyleRun untrusted JavaScript and full dev servers, with git hosting and domains built in.65.5

Runtime

Cold start (ms)
700ms
Cold start (claimed) (ms)
Max runtime (ms)
60min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
No
GPU
No
Runtimes
javascript-typescriptpython
Preview URLs
Yes
Browser inside
No
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
·

Isolation & network

Isolation
microvm
Root in guest
Partial
Egress control
allowlist

Agent integration

MCP server
Partial
SDKs
typescriptrest-api
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-request
Idle cost
storage-only

Traction

GitHub stars
·
Funding
seed

Positioning

Positioning
The backend for AI app builders: run code, host git, serve previews, attach domains.

Timeline

Shipped
2024
exe.dev logoexe.devPersistent VMs you SSH into, with root, apt and systemd, on a flat monthly plan.64.9

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Unknown
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Unknown
Browser inside
Unknown
File up/download
Partial
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost

Traction

GitHub stars
·
Funding
·

Positioning

Positioning
Modern VMs you SSH into, with root, apt and systemd, on a flat monthly plan.

Timeline

Shipped
·
ascii logoasciiAgent orchestration over Telegram, running on box's VM infrastructure.64.7

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Partial
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Partial
Browser inside
Unknown
File up/download
Unknown
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost

Traction

GitHub stars
·
Funding
·

Positioning

Positioning
A pocket CTO on Telegram: multi-repo dispatch and cross-VM agent comms, running on box.

Timeline

Shipped
·
box logoboxPersistent Linux VMs with SSH, per-VM IPv4 and disk-level forking, priced flat.64.2

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Partial
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Partial
Browser inside
Unknown
File up/download
Partial
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost

Traction

GitHub stars
·
Funding
·

Positioning

Positioning
A persistent Linux VM with SSH, disk-level forking, a dedicated IPv4 and Docker inside.

Timeline

Shipped
·
GitHub Codespaces logoGitHub CodespacesDevcontainer-backed cloud VMs built for humans, occasionally repurposed for agents.62.8

Runtime

Cold start (ms)
45s
Cold start (claimed) (ms)
10s
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
No
Runtimes
any-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Partial
Sydney region
Yes
Concurrent limit (sandboxes)
10sandboxes

Isolation & network

Isolation
vm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
No
SDKs
rest-apicli
Streaming output
Partial

Pricing

Price / hour (/hr)
$0.18/hr
Meter
per-minute
Idle cost
storage-only

Traction

GitHub stars
·
Funding
public

Positioning

Positioning
A configured development environment in the cloud, one click from the repository.

Timeline

Shipped
2020
Northflank logoNorthflankContainer platform with bring-your-own-cloud, GPUs and a sandbox-shaped API.62.6

Runtime

Cold start (ms)
12s
Cold start (claimed) (ms)
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
No
Custom images
Yes
GPU
Yes
Runtimes
any-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Partial
Sydney region
Partial
Concurrent limit (sandboxes)
200sandboxes

Isolation & network

Isolation
container
Root in guest
Partial
Egress control
allowlist

Agent integration

MCP server
Yes
SDKs
typescriptrest-api+1
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.07/hr
Meter
per-minute
Idle cost
full-rate

Traction

GitHub stars
·
Funding
series-a

Positioning

Positioning
Run agent workloads, services and GPUs in your own cloud account.

Timeline

Shipped
2021
Self-hosted Firecracker logoSelf-hosted FirecrackerThe baseline: Firecracker on your own metal, plus every hard part you now own.61.2

Runtime

Cold start (ms)
200ms
Cold start (claimed) (ms)
125ms
Max runtime (ms)
1440min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
Partial
Runtimes
any-oci-image
Preview URLs
No
Browser inside
Partial
File up/download
No
Sydney region
Yes
Concurrent limit (sandboxes)
1,000sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
No
SDKs
rest-api
Streaming output
No

Pricing

Price / hour (/hr)
$0.03/hr
Meter
per-hour
Idle cost
full-rate

Traction

GitHub stars
27,000
Funding
·

Positioning

Positioning
The microVM monitor underneath most of this roster, and none of the platform.

Timeline

Shipped
2018
Cloudflare Sandbox logoCloudflare SandboxContainer sandboxes driven from a Worker, addressed through Durable Objects.60.1

Runtime

Cold start (ms)
3s
Cold start (claimed) (ms)
Max runtime (ms)
60min
Persistent FS
No
Snapshot & fork
No
Custom images
Yes
GPU
No
Runtimes
any-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
Partial
Concurrent limit (sandboxes)
100sandboxes

Isolation & network

Isolation
container
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
Partial
SDKs
typescript
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.18/hr
Meter
per-10ms
Idle cost
free-when-stopped

Traction

GitHub stars
·
Funding
public

Positioning

Positioning
Containers you can call from a Worker, with a Durable Object as the control plane.

Timeline

Shipped
2025
Val Town logoVal TownDeno isolates that run TypeScript on HTTP, cron and email triggers.44.8

Runtime

Cold start (ms)
150ms
Cold start (claimed) (ms)
·
Max runtime (ms)
5min
Persistent FS
No
Snapshot & fork
No
Custom images
No
GPU
No
Runtimes
javascript-typescript
Preview URLs
Yes
Browser inside
No
File up/download
Partial
Sydney region
No
Concurrent limit (sandboxes)

Isolation & network

Isolation
v8-isolate
Root in guest
No
Egress control
proxy-logged

Agent integration

MCP server
Yes
SDKs
typescriptrest-api+1
Streaming output
Partial

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost
free-when-stopped

Traction

GitHub stars
·
Funding
seed

Positioning

Positioning
Write and deploy TypeScript in the browser, triggered by HTTP, cron or email.

Timeline

Shipped
2022
Riza logoRizaWASM-isolated code interpreter API for LLM output, with no filesystem or network.42.9

Runtime

Cold start (ms)
120ms
Cold start (claimed) (ms)
Max runtime (ms)
1min
Persistent FS
No
Snapshot & fork
No
Custom images
No
GPU
No
Runtimes
pythonjavascript-typescript+2
Preview URLs
No
Browser inside
No
File up/download
Partial
Sydney region
No
Concurrent limit (sandboxes)

Isolation & network

Isolation
wasm
Root in guest
No
Egress control
blocked-by-default

Agent integration

MCP server
Partial
SDKs
pythontypescript+1
Streaming output
No

Pricing

Price / hour (/hr)
Meter
per-request
Idle cost
free-when-stopped

Traction

GitHub stars
·
Funding
seed

Positioning

Positioning
A code interpreter API for safely running code written by LLMs.

Timeline

Shipped
2024
yespartialnounknown
Sources shown beside each value · Learn how sourcing works
measured vendor-claimed community inferredExpand any row for the source, verification date, and caveat behind every cell.

Where each one wins

  • You need Australian or multi-region placement
  • Cost per sandbox-hour matters more than SDK ergonomics
  • You already run on Fly and can reuse the networking

Unbeatable price-per-isolation and the only mainstream Sydney region, paid for in a week of glue code.

Blaxel

79.6
  • You want sandboxes, agent hosting and MCP from one vendor
  • Latency-sensitive workloads where you can keep a pool warm
  • You are comfortable betting on an early-stage platform

The 25 ms claim measures the wrong operation, and the broad platform is a lot to buy from a seed-stage company.

  • Boot time is your headline metric and the code is your own
  • You want an MCP server maintained by the vendor
  • You can tolerate a fast-moving, now closed-source product surface

E2B's most credible ergonomics challenger with the best MCP story, but shared-kernel containers make it wrong for untrusted code.

  • Running SWE-bench-style evaluation loops at scale
  • Generating training data from reproducible environments
  • You want a vendor whose roadmap targets agent researchers

Built around evaluation, not production traffic: the obvious pick if you measure an agent's pass rate nightly.

Modal

77.5
  • The sandbox needs a GPU on a self-serve plan
  • Your team is Python-first and already writes Modal functions
  • One platform for batch ML and untrusted execution

Buy Modal for GPUs and Python ergonomics, not boot time; unmatched when the agent runs inference next to its code.

  • An app-builder or code-runner inside a Vercel deployment
  • Agent sessions that mostly wait on a model, where active-CPU billing pays off
  • You need a per-sandbox egress allowlist without building it yourself

Now a broad, well-isolated product with active-CPU billing, held back only by a single US region.

E2B

74.4
  • You want the category default with a real self-host escape hatch
  • Your environments are already Dockerfiles
  • Python code-interpreter workloads at steady volume

The safest default: complete SDKs, Dockerfile templates and an Apache licence, though no longer fastest and weak at hobby scale.

  • You fork one prepared environment many times
  • Long-lived per-user environments that hibernate between sessions
  • You need processes and memory preserved, not just the disk

The best-engineered snapshot-and-fork here for branching workloads, now a component of Together AI rather than its own product.

  • You already use Namespace for CI runners
  • Build-cache-heavy environments that are slow to reconstruct
  • You want infrastructure primitives rather than an agent SDK

Snapshotted microVMs with good caching but no agent SDK, worth a look mainly if your CI already runs here.

Islo

65.8
  • Central policy control over what each agent can reach
  • Enterprise teams that must govern agent identity and access
  • You are already inside Incredibuild's SDLC platform

An enterprise governance pitch this rival's table cannot capture; verify from Incredibuild's own documentation before shortlisting.

  • Building an AI app builder with live previews
  • You want git hosting and custom domains bundled with execution
  • JavaScript and TypeScript workloads only

Collapses four vendors into one for Lovable-shaped app builders, wrong shape for anything outside JavaScript.

  • Long-lived SSH VMs on a predictable monthly bill
  • You want root, apt and systemd rather than a constrained runtime
  • You are prepared to verify the gaps a competitor's chart reports

A placeholder built from a rival's table: positive claims credited, omissions ignored, nothing measured.

ascii

64.7
  • You want an agent you message rather than an API you call
  • Multi-repo dispatch across long-running VMs
  • You are already using box and want the harness on top

An agent harness with a VM attached, judged from an unverified vendor table with no fundamentals established.

box

64.2
  • You want a VM that stays up rather than a sandbox that vanishes
  • SSH, Docker and a real network stack matter more than an SDK
  • You are willing to test the vendor's claims yourself before committing

A coherent long-lived-VM pitch, but a vendor self-assessment with no measured isolation, boot time or price.

  • Human-in-the-loop agent sessions rather than autonomous fleets
  • You need an Australian region and are already on GitHub
  • Your environments are already defined as devcontainers

Excellent for a human, a bad substrate for an API creating hundreds of sandboxes.

  • Compliance requires the workload inside your own VPC
  • Sandboxes are long-lived per customer, not per request
  • You also need databases, jobs and GPUs on one platform

Loses on boot time, wins on running inside your own cloud account and compliance boundary.

  • Sustained volume where vendor margin exceeds an engineer's salary
  • Hard data-residency or air-gap requirements
  • You already operate bare metal and a scheduler

The honest baseline: identical isolation at a tenth the cost, minus the two engineer-years of orchestration.

  • Your app already runs on Workers and Durable Objects
  • You want code-level control over what the sandbox can reach
  • Sandboxes are long-lived per user rather than created in bursts

An addressable sandbox object in your app, worth it on Workers but too slow to migrate to.

  • Evaluating short TypeScript snippets behind an HTTP endpoint
  • Cron and webhook glue an agent can write and deploy itself
  • Zero infrastructure, and you can accept zero filesystem

The isolate benchmark: fastest and cheapest for a TypeScript snippet behind a URL, outgrown the moment you need `npm install`.

Riza

42.9
  • Executing model-written expressions with the tightest possible default
  • Multi-language interpreter tooling without maintaining images
  • You want deny-by-default network and filesystem out of the box

The strictest isolation here: perfect for evaluating a model's data transform, useless if it needs a dependency.

Which one should you pick?

The short path to an answer. Follow the branch that matches your situation; the table above is there when you want to check the reasoning against the numbers.
What are you actually running in the sandbox?
  • Agent-generated code, and I want a default that just works
    Do you need to self-host, or keep an exit hatch if the vendor changes?
    • Yes — portability matters
      E2B logoE2BApache-licensed Firecracker with Dockerfile templates; the only ecosystem you can self-host without rewriting your integration.
    • No — fully managed is fine, and boot time is the whole game
      Daytona logoDaytonaSub-90 ms shared-kernel container start wins on ergonomics and speed — but a shared kernel is not a hypervisor, so weigh the isolation you are giving up.
  • Evals and benchmark-harness work
    Runloop logoRunloopDevbox snapshots plus a built-in benchmark harness — a different product shape that fits evals rather than general execution.
  • A feature inside an app I've already deployed
    Where does that app already live?
    • Vercel
      Vercel Sandbox logoVercel SandboxMakes the sandbox a line in the app you already ship; not fastest or cheapest, and not trying to be.
    • Cloudflare
      Cloudflare Sandbox logoCloudflare SandboxThe same bet on Workers: the sandbox is part of the deployment you already run.
    • My own infra
      Fly.io Machines logoFly.io MachinesA real Firecracker VM at roughly a tenth the cost, with a durable volume, a Sydney region, and nothing charged while stopped.

Which sandbox provider should I run untrusted or agent-generated code on?

  • Nobody buys a sandbox; they buy the seconds between an agent deciding to run code and it running
  • They also buy confidence that a hostile rm -rf or crypto miner stays inside the box
  • Everything else on this page is a proxy for those two things
  • Fastest: V8 isolates and WASM (Val TownVal Town logo, RizaRiza logo) start in milliseconds
  • Mid: snapshot microVMs (E2BE2B logo, BlaxelBlaxel logo, CodeSandbox SDKCodeSandbox SDK logo, RunloopRunloop logo) land between 150 ms and two seconds
  • Slowest: Kubernetes containers (NorthflankNorthflank logo) and full VMs (GitHub CodespacesGitHub Codespaces logo), seconds to tens of seconds
  • DaytonaDaytona logo's sub-90 ms is a shared-kernel container start, no guest kernel to boot
  • BlaxelBlaxel logo's 25 ms and DaytonaDaytona logo's sub-90 ms are real, but only for warm operations on your own machine
  • Not the number you get over TLS into a pool that just scaled to zero
  • E2BE2B logo is the default for agent harnesses: Apache-licensed, Firecracker, Dockerfile templates
  • E2BE2B logo is the only ecosystem where you can self-host without rewriting your integration
  • DaytonaDaytona logo wins on ergonomics and boot time, loses on isolation: a shared kernel is not a hypervisor
  • DaytonaDaytona logo closed its source in June 2026, removing the exit hatch that makes E2BE2B logo's licence matter
  • RunloopRunloop logo suits evals: devbox snapshots plus a benchmark harness is a different product shape
  • BlaxelBlaxel logo and FreestyleFreestyle logo bet the company on boot time, a fine bet and a thin moat
  • If you already have infra, Fly.io MachinesFly.io Machines logo gives a real Firecracker VM for ~a tenth the cost
  • Fly adds a durable volume, public IPv4, Sydney region, per-second billing, nothing charged while stopped
  • Vercel SandboxVercel Sandbox logo and Cloudflare SandboxCloudflare Sandbox logo make the sandbox a line in the app you already deployed
  • Neither VercelVercel logo nor Cloudflare is fastest or cheapest, and neither is trying to be
  • Snapshot-and-fork is becoming table stakes, killing cold-start as a standalone pitch
  • Once everyone restores a paused VM in 200 ms, the differentiator moves to network reach
  • Expect egress control, per-sandbox audit logs and outbound proxies to be the 2027 argument
  • Enterprise buyers block deals on egress, so Fly, NorthflankNorthflank logo, Cloudflare and self-hosted Firecracker are well placed

What cold start actually means

  • At least five distinct events get called "cold start"
  • Snapshot restore: a booted VM's memory image mapped back on a host that has the pages, the source of sub-100 ms numbers
  • Snapshot restore presumes hardware allocated, image cached and pool warm
  • Guest boot: Firecracker to userspace, ~125 ms, where the microVM reputation comes from
  • Scheduling: finding a host with capacity, zero when warm, seconds when not
  • Image pull: zero for a cached base, thirty seconds for your 4 GB CUDA layer
  • API round trip and TLS: 20 ms from the same region, 300 ms from Sydney
  • Vendors quote the smallest of these; agents pay all five
  • The measured column starts from a client with nothing running and stops when a command returns output
  • Within 3× of the claim is good: E2BE2B logo 2.0×, Modal 1.5×, CodeSandbox 1.2×, Fly 3.0×, DaytonaDaytona logo 3.9×
  • BlaxelBlaxel logo is an order of magnitude out: 25 ms snapshot restore vs ~20× estimated create-to-output
  • Gap size tells you how a vendor defines the word, not how fast its infrastructure is

The pricing traps

  • Hourly compute is compared most and rarely dominates the bill
  • Watch the meter's units: per-second vs per-minute on 400 twelve-second sandboxes is a completely different bill
  • Per-minute rounds every one of them up
  • Cloudflare's 10 ms granularity and VercelVercel logo's active-CPU model suit bursty agent traffic; minute blocks punish it
  • Watch idle: an agent session is mostly the sandbox waiting for a model to generate
  • Wall-clock billing means renting a CPU to wait on someone else's inference queue
  • Active-CPU or pause-to-storage vendors can be several times cheaper despite a higher headline rate
  • Check the pause API's latency first: pausing something you resume 400 ms later is a false economy
  • Hidden line items: snapshot storage per GB-month never garbage-collected, egress at cloud rates
  • More hidden costs: volumes that outlive their sandbox, and platform minimums like a paid Workers plan
  • A "$0.05/hr" provider can cost $150 before the first sandbox boots

How to choose in an afternoon

  • Start with isolation, the only requirement you cannot retrofit
  • Untrusted code needs a kernel boundary (microVM or gVisor) and outbound network control
  • That shortlists Fly, NorthflankNorthflank logo, Cloudflare, E2BE2B logo and self-hosted Firecracker
  • Your own agent's code on your own data is a lower bar every provider clears
  • Then decide: buying infrastructure or buying a product
  • If the sandbox is your core loop, take the agent-native SDK and snapshot semantics (E2BE2B logo, DaytonaDaytona logo, RunloopRunloop logo)
  • Treat the per-hour premium as the price of not maintaining an orchestrator
  • If it is one feature, use what your platform gives you and revisit when the bill or boot time hurts
  • Run one benchmark before signing: your real image, your real region, 50 cold creates, record p95 not p50
  • p50 is a good day; p95 is what users see when the pool is cold, deciding instant vs broken
  • If a vendor won't let you run that test on a trial account, that is the answer

What is sandbox providers compared: cold start, isolation and price?

Sandbox providers compared: cold start, isolation and price

On toolweight, Sandbox providers compared: cold start, isolation and price means the 20 tools benchmarked on this page, E2BE2B logo, Modal, DaytonaDaytona logo, Fly.io MachinesFly.io Machines logo, Cloudflare SandboxCloudflare Sandbox logo, Vercel SandboxVercel Sandbox logo, CodeSandbox SDKCodeSandbox SDK logo, NorthflankNorthflank logo, RunloopRunloop logo, BlaxelBlaxel logo, FreestyleFreestyle logo, Val TownVal Town logo, RizaRiza logo, Namespace, GitHub CodespacesGitHub Codespaces logo, box, ascii, exe.devexe.dev logo, IsloIslo logo, Self-hosted FirecrackerSelf-hosted Firecracker logo, judged on the same 26 fields, from the same sources, on the same date. The question it exists to answer: Which sandbox provider should I run untrusted or agent-generated code on?

How does toolweight compare these?

  • Field values come from vendor docs, pricing pages and changelogs, each cell carrying its provenance
  • Prices are list rate for ~2 vCPU and 4 GB, no commitment, normalised to hourly
  • Per-second or credit vendors are marked inferred with arithmetic noted, or left null
  • Claimed cold start is whatever the vendor puts on its homepage, scored at weight zero
  • Vendors quote wildly different things: snapshot restore, guest boot, API return, or warm-pool p50
  • Measured cold start is toolweight's own harness: single client, fixed region, TLS included
  • It calls create-sandbox and blocks until a trivial command returns output, the full round trip
  • Runs on an account with nothing alive after an idle period, so nothing is pool-warm
  • Reports p50 and p95 over at least 50 runs so one good afternoon cannot flatter the number
  • No cell currently carries measured confidence; the column is inferred estimates for now
  • We would rather show an honest estimate than launder a vendor's number into a benchmark
  • Four rows (box, ascii, exe.devexe.dev logo, IsloIslo logo) are transcribed from box's own table at box.ascii.dev/compare
  • That is the least neutral source here: a vendor's chart, its products first, rivals arranged around them
  • toolweight has verified none of it; fields the table omits are left null and unknown
  • Only a cell the table states directly is marked vendor-claimed
  • Derived values (root off "Docker inside the VM", preview-URL off IP rows) are marked inferred
  • A marketing row is not a measurement: "runs 24/7" and "1000+ concurrent VMs" stay unknown
  • Absence is scored asymmetrically: a vendor omitting its own product is credited
  • box and ascii are absent from the process-fork and sub-500 ms rows, and that is recorded
  • A vendor omitting a rival counts as nothing: exe.devexe.dev logo and IsloIslo logo's snapshot cells stay unknown
  • These four rows stay this way until re-sourced from each vendor's own documentation
Full methodology and sourcing policy →

What has changed on this page

Every value carries a source and a date; when one moves, it is logged here. Spot a number that is wrong or stale? Every cell in the table has a "dispute" link that opens a prefilled correction.
  1. 2026-07-23Changed
    Held every measured cold-start as inferred, not measured: the harness has not run on this roster yet, and an honest estimate beats laundering a vendor's homepage figure into a benchmark. The latency probe backfills this column as it runs.
    Cold start
  2. 2026-06-30Corrected
    Recorded Daytona closing its source in June 2026, which removes the self-host exit hatch that made its shared-kernel isolation trade-off tolerable.
  3. 2026-06-15Added
    Added box, ascii, exe.dev and Islo, transcribed from box's own comparison table; every unverified cell is left unknown until re-sourced from each vendor's own docs.
Every correction, as public issues
Cite this comparisonCC-BY-4.0 · verified 2026-07-23

E2B is the fastest general-purpose microVM sandbox for agents, but Fly.io Machines wins on price and control if you can skip an agent SDK. — toolweight, https://toolweight.com/compare/sandbox-providers, verified 2026-07-23. Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

Frequently asked questions

Is a container enough to run LLM-generated code, or do I need a microVM?

  • Your own model's code behind your own prompt: a hardened container is usually fine
  • User-supplied code: assume a container escape is a matter of time and budget, shared-kernel isolation drips CVEs
  • A microVM (Firecracker, Cloud Hypervisor) or gVisor gives a syscall boundary a guest kernel bug cannot cross
  • Check what you buy: DaytonaDaytona logo runs shared-kernel containers, and vendors call a namespace a "dedicated kernel"
  • The price gap is now small enough that the container answer is rarely worth defending

Why is the cold start I measure so much slower than the number on the vendor's homepage?

  • They measure different things
  • Vendor numbers time a snapshot restore on hardware already allocated to you
  • Yours adds TLS, API scheduling, image pull, the round trip, and pool scale-up queueing
  • A 90 ms claim routinely lands between 400 ms and 2 s, and widens the further you sit from their region

What does snapshot and fork actually buy an agent?

  • Branching: build once, snapshot, then fork five times to try five patches in parallel
  • Cheap idle: pause mid-session, pay storage only, resume in a fraction of a cold boot
  • Resume keeps the process tree and page cache intact
  • For long agent sessions this beats raw boot speed, because you cold-start once

Can I run a browser inside the sandbox instead of paying for a browser API?

  • Yes, and for a handful of pages a day it is cheaper
  • Chromium with a CDP endpoint costs compute you already pay for
  • You lose residential egress, CAPTCHA handling and fingerprint work, so real bot defences will beat you
  • Use the sandbox for your own apps and tooling; use a browser API when the site fights back

How locked in am I if I pick the wrong provider?

  • Less than it feels: every provider exposes the same four verbs, create, exec, read/write files, kill
  • A thin interface of your own keeps swapping to about a day of work
  • Real lock-in is proprietary image formats, snapshots that only restore on the vendor's cluster, and baked-in preview-URL domains
  • Keep images as plain Dockerfiles and you keep your exit

Which providers can run in Sydney or another Australian region?

Do I pay while a sandbox sits idle waiting for the model to respond?

  • It depends on the meter, and this is where bills go wrong
  • Fly charges nothing for a stopped machine beyond storage
  • E2BE2B logo and CodeSandbox bill storage on a paused sandbox; VercelVercel logo bills active CPU, so network-blocked costs little
  • Wall-clock per-second vendors charge for every token the model is still generating
  • On agent workloads that idle time is usually the majority of the session