Sandbox providers

E2B is the fastest general-purpose microVM sandbox for agents; Daytona is quicker still but runs containers, not microVMs, so it is the wrong default for untrusted code. Fly.io Machines wins on control, regions and price if you can live without an agent SDK. Vercel Sandbox and Cloudflare Sandbox are easiest if you already deploy there. Every published cold-start number is marketing until someone measures it cold.

Last verified
2026-07-23 (today)
Source confidence
32%
Re-verified
every 30 days
Tools
20
Fields
26
20 tools · verified today
Fly.io MachinesRaw Firecracker microVMs with a REST API, durable volumes and 35+ regions.80.8

Runtime

Cold start (ms)
900 ms
Cold start (claimed) (ms)
300 ms
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
Yes
Runtimes
any-oci-image
Preview URLs
Partial
Browser inside
Partial
File up/download
Partial
Sydney region
Yes
Concurrent limit (sandboxes)
500 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
Partial
SDKs
rest-apicli
Streaming output
Partial

Pricing

Price / hour (/hr)
$0.031 /hr
Meter
per-second
Idle cost
free-when-stopped

Traction

GitHub stars
Funding
series-c-plus

Positioning

Positioning
Hardware-virtualised containers that run anywhere, started and stopped over an API.

Timeline

Shipped
2022
BlaxelAgent-first cloud claiming ~25 ms microVM boots from snapshots.79.6

Runtime

Cold start (ms)
500 ms
Cold start (claimed) (ms)
25 ms
Max runtime (ms)
1440 min
Persistent FS
Partial
Snapshot & fork
Yes
Custom images
Yes
GPU
Unknown
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
Partial
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
Yes
SDKs
pythontypescriptcli
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
Funding
seed

Positioning

Positioning
An agent-native cloud: sandboxes, agent hosting and a model gateway in one platform.

Timeline

Shipped
2025
DaytonaSub-second container sandboxes for agent workloads, from a team that built a dev-env manager.79.4

Runtime

Cold start (ms)
350 ms
Cold start (claimed) (ms)
90 ms
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
Unknown
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
container
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
Yes
SDKs
pythontypescriptcli
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.17 /hr
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
Funding
seed

Positioning

Positioning
Secure, elastic infrastructure for running AI-generated code.

Timeline

Shipped
2025
RunloopDevboxes, blueprints and eval harnesses aimed at teams training coding agents.79.0

Runtime

Cold start (ms)
800 ms
Cold start (claimed) (ms)
1 s
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
Partial
SDKs
pythontypescriptrest-apicli
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
Funding
seed

Positioning

Positioning
Infrastructure and evaluation harnesses for teams building coding agents.

Timeline

Shipped
2024
ModalServerless Python compute with a Sandbox API bolted onto a GPU-first platform.77.5

Runtime

Cold start (ms)
1.5 s
Cold start (claimed) (ms)
1 s
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
Yes
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
gvisor
Root in guest
Yes
Egress control
proxy-logged

Agent integration

MCP server
Partial
SDKs
pythoncli
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.238 /hr
Meter
per-second
Idle cost
active-cpu-only

Traction

GitHub stars
Funding
series-b

Positioning

Positioning
Serverless compute for AI teams, with sandboxes as one primitive.

Timeline

Shipped
2024
Vercel SandboxFirecracker microVMs for running untrusted code from inside a Vercel deployment.74.6

Runtime

Cold start (ms)
1.8 s
Cold start (claimed) (ms)
Max runtime (ms)
1440 min
Persistent FS
Partial
Snapshot & fork
Partial
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
No
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
2,000 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
No
SDKs
typescript
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.341 /hr
Meter
per-second
Idle cost
active-cpu-only

Traction

GitHub stars
Funding
series-c-plus

Positioning

Positioning
Run untrusted code from your Vercel app in an ephemeral microVM.

Timeline

Shipped
2025
CodeSandbox SDKFirecracker VMs with memory snapshots, from the online IDE, now owned by Together AI.74.4

Runtime

Cold start (ms)
1.2 s
Cold start (claimed) (ms)
1 s
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
No
SDKs
typescriptpython
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
13,000
Funding
acquired

Positioning

Positioning
The VM infrastructure behind CodeSandbox, sold as an SDK.

Timeline

Shipped
2024
E2BOpen-source Firecracker sandboxes with Python and TypeScript SDKs for AI agents.74.4

Runtime

Cold start (ms)
400 ms
Cold start (claimed) (ms)
200 ms
Max runtime (ms)
1440 min
Persistent FS
Partial
Snapshot & fork
Partial
Custom images
Yes
GPU
No
Runtimes
pythonjavascript-typescriptbashany-oci-image
Preview URLs
Yes
Browser inside
Yes
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
open

Agent integration

MCP server
Partial
SDKs
pythontypescriptcli
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.17 /hr
Meter
per-second
Idle cost
storage-only

Traction

GitHub stars
9,000
Funding
series-a

Positioning

Positioning
Open-source infrastructure for running AI-generated code.

Timeline

Shipped
2023
NamespaceFast microVM instances and CI runners with instant snapshots and cached builds.72.6

Runtime

Cold start (ms)
1.5 s
Cold start (claimed) (ms)
800 ms
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Yes
Custom images
Yes
GPU
Unknown
Runtimes
any-oci-image
Preview URLs
Partial
Browser inside
Partial
File up/download
Partial
Sydney region
No
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
proxy-logged

Agent integration

MCP server
No
SDKs
gorest-apicli
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-minute
Idle cost
free-when-stopped

Traction

GitHub stars
Funding
seed

Positioning

Positioning
Fast ephemeral compute for CI, builds and remote development.

Timeline

Shipped
2023
IsloPer-agent isolated cloud sandboxes with enterprise policy controls, from Incredibuild.65.8

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Unknown
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Partial
Browser inside
Unknown
File up/download
Partial
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
Idle cost

Traction

GitHub stars
Funding

Positioning

Positioning
An isolated cloud environment per coding agent, with enterprise policy control over what it can reach.

Timeline

Shipped
FreestyleRun untrusted JavaScript and full dev servers, with git hosting and domains built in.65.5

Runtime

Cold start (ms)
700 ms
Cold start (claimed) (ms)
Max runtime (ms)
60 min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
No
GPU
No
Runtimes
javascript-typescriptpython
Preview URLs
Yes
Browser inside
No
File up/download
Yes
Sydney region
No
Concurrent limit (sandboxes)

Isolation & network

Isolation
microvm
Root in guest
Partial
Egress control
allowlist

Agent integration

MCP server
Partial
SDKs
typescriptrest-api
Streaming output
Yes

Pricing

Price / hour (/hr)
Meter
per-request
Idle cost
storage-only

Traction

GitHub stars
Funding
seed

Positioning

Positioning
The backend for AI app builders: run code, host git, serve previews, attach domains.

Timeline

Shipped
2024
exe.devPersistent VMs you SSH into, with root, apt and systemd, on a flat monthly plan.64.9

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Unknown
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Unknown
Browser inside
Unknown
File up/download
Partial
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost

Traction

GitHub stars
Funding

Positioning

Positioning
Modern VMs you SSH into, with root, apt and systemd, on a flat monthly plan.

Timeline

Shipped
asciiAgent orchestration over Telegram, running on box's VM infrastructure.64.7

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Partial
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Partial
Browser inside
Unknown
File up/download
Unknown
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost

Traction

GitHub stars
Funding

Positioning

Positioning
A pocket CTO on Telegram: multi-repo dispatch and cross-VM agent comms, running on box.

Timeline

Shipped
boxPersistent Linux VMs with SSH, per-VM IPv4 and disk-level forking, priced flat.64.2

Runtime

Cold start (ms)
Cold start (claimed) (ms)
Max runtime (ms)
Persistent FS
Unknown
Snapshot & fork
Partial
Custom images
Partial
GPU
Unknown
Runtimes
Preview URLs
Partial
Browser inside
Unknown
File up/download
Partial
Sydney region
Unknown
Concurrent limit (sandboxes)

Isolation & network

Isolation
Root in guest
Yes
Egress control

Agent integration

MCP server
Unknown
SDKs
Streaming output
Unknown

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost

Traction

GitHub stars
Funding

Positioning

Positioning
A persistent Linux VM with SSH, disk-level forking, a dedicated IPv4 and Docker inside.

Timeline

Shipped
GitHub CodespacesDevcontainer-backed cloud VMs built for humans, occasionally repurposed for agents.62.8

Runtime

Cold start (ms)
45 s
Cold start (claimed) (ms)
10 s
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
No
Runtimes
any-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Partial
Sydney region
Yes
Concurrent limit (sandboxes)
10 sandboxes

Isolation & network

Isolation
vm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
No
SDKs
rest-apicli
Streaming output
Partial

Pricing

Price / hour (/hr)
$0.18 /hr
Meter
per-minute
Idle cost
storage-only

Traction

GitHub stars
Funding
public

Positioning

Positioning
A configured development environment in the cloud, one click from the repository.

Timeline

Shipped
2020
NorthflankContainer platform with bring-your-own-cloud, GPUs and a sandbox-shaped API.62.6

Runtime

Cold start (ms)
12 s
Cold start (claimed) (ms)
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
No
Custom images
Yes
GPU
Yes
Runtimes
any-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Partial
Sydney region
Partial
Concurrent limit (sandboxes)
200 sandboxes

Isolation & network

Isolation
container
Root in guest
Partial
Egress control
allowlist

Agent integration

MCP server
Yes
SDKs
typescriptrest-apicli
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.07 /hr
Meter
per-minute
Idle cost
full-rate

Traction

GitHub stars
Funding
series-a

Positioning

Positioning
Run agent workloads, services and GPUs in your own cloud account.

Timeline

Shipped
2021
Self-hosted FirecrackerThe baseline: Firecracker on your own metal, plus every hard part you now own.61.2

Runtime

Cold start (ms)
200 ms
Cold start (claimed) (ms)
125 ms
Max runtime (ms)
1440 min
Persistent FS
Yes
Snapshot & fork
Partial
Custom images
Yes
GPU
Partial
Runtimes
any-oci-image
Preview URLs
No
Browser inside
Partial
File up/download
No
Sydney region
Yes
Concurrent limit (sandboxes)
1,000 sandboxes

Isolation & network

Isolation
microvm
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
No
SDKs
rest-api
Streaming output
No

Pricing

Price / hour (/hr)
$0.03 /hr
Meter
per-hour
Idle cost
full-rate

Traction

GitHub stars
27,000
Funding

Positioning

Positioning
The microVM monitor underneath most of this roster, and none of the platform.

Timeline

Shipped
2018
Cloudflare SandboxContainer sandboxes driven from a Worker, addressed through Durable Objects.60.1

Runtime

Cold start (ms)
3 s
Cold start (claimed) (ms)
Max runtime (ms)
60 min
Persistent FS
No
Snapshot & fork
No
Custom images
Yes
GPU
No
Runtimes
any-oci-image
Preview URLs
Yes
Browser inside
Partial
File up/download
Yes
Sydney region
Partial
Concurrent limit (sandboxes)
100 sandboxes

Isolation & network

Isolation
container
Root in guest
Yes
Egress control
allowlist

Agent integration

MCP server
Partial
SDKs
typescript
Streaming output
Yes

Pricing

Price / hour (/hr)
$0.18 /hr
Meter
per-10ms
Idle cost
free-when-stopped

Traction

GitHub stars
Funding
public

Positioning

Positioning
Containers you can call from a Worker, with a Durable Object as the control plane.

Timeline

Shipped
2025
Val TownDeno isolates that run TypeScript on HTTP, cron and email triggers.44.8

Runtime

Cold start (ms)
150 ms
Cold start (claimed) (ms)
Max runtime (ms)
5 min
Persistent FS
No
Snapshot & fork
No
Custom images
No
GPU
No
Runtimes
javascript-typescript
Preview URLs
Yes
Browser inside
No
File up/download
Partial
Sydney region
No
Concurrent limit (sandboxes)

Isolation & network

Isolation
v8-isolate
Root in guest
No
Egress control
proxy-logged

Agent integration

MCP server
Yes
SDKs
typescriptrest-apicli
Streaming output
Partial

Pricing

Price / hour (/hr)
Meter
subscription
Idle cost
free-when-stopped

Traction

GitHub stars
Funding
seed

Positioning

Positioning
Write and deploy TypeScript in the browser, triggered by HTTP, cron or email.

Timeline

Shipped
2022
RizaWASM-isolated code interpreter API for LLM output, with no filesystem or network.42.9

Runtime

Cold start (ms)
120 ms
Cold start (claimed) (ms)
Max runtime (ms)
1 min
Persistent FS
No
Snapshot & fork
No
Custom images
No
GPU
No
Runtimes
pythonjavascript-typescriptrubyphp
Preview URLs
No
Browser inside
No
File up/download
Partial
Sydney region
No
Concurrent limit (sandboxes)

Isolation & network

Isolation
wasm
Root in guest
No
Egress control
blocked-by-default

Agent integration

MCP server
Partial
SDKs
pythontypescriptrest-api
Streaming output
No

Pricing

Price / hour (/hr)
Meter
per-request
Idle cost
free-when-stopped

Traction

GitHub stars
Funding
seed

Positioning

Positioning
A code interpreter API for safely running code written by LLMs.

Timeline

Shipped
2024
yespartialnounknown◆ measured · ▸ vendor-claimed · ▪ community · · inferredExpand a row for the source and date behind every cell.
Sandbox providers compared: cold start, isolation and price pricing, compared line by line

Which sandbox provider should I run untrusted or agent-generated code on?

Nobody buys a sandbox. They buy the two seconds between an agent deciding to run code and the code running, and they buy the confidence that a hostile rm -rf or a crypto miner stays inside the box. Everything else on this page is a proxy for those two things.

On speed, the honest ranking is: V8 isolates and WASM (Val Town, Riza) start in milliseconds because they barely start anything; microVM providers with memory snapshots (E2B, Blaxel, CodeSandbox SDK, Runloop) land somewhere between 150 ms and two seconds depending on how warm their pool is when you call; container platforms scheduled on Kubernetes (Northflank) and full VMs (GitHub Codespaces) are seconds to tens of seconds. Daytona sits outside that grouping and it matters: its sandboxes are Docker/OCI containers on a shared host kernel, which is precisely why it can claim a create in under 90 ms — there is no guest kernel to boot. Blaxel's 25 ms and Daytona's sub-90 ms are real numbers for real operations, restoring a snapshot or starting a container on a machine that is already yours. They are not the number you get from your laptop, over TLS, into a pool that has just scaled to zero. That gap is what the killer column on this page exists to expose.

For agent harnesses specifically, E2B remains the default and deserves it: Apache-licensed, Firecracker underneath, a template system that is just a Dockerfile, and the only ecosystem where you can rip the vendor out and self-host without rewriting your integration. Daytona is the sharpest competitor on ergonomics and boot time, and the weakest on the one axis this page treats as non-negotiable — a shared kernel is a different security posture from a hypervisor, whatever the marketing says about a "dedicated kernel". It closed its source in June 2026, too, which removes the exit hatch that makes E2B's licence worth paying attention to. Runloop is the one to look at if you are running evals rather than a product — devbox snapshots plus a benchmark harness is a genuinely different shape of product. Blaxel and Freestyle are betting the whole company on boot time, which is a fine bet and a thin moat.

If you already have infrastructure, the calculus inverts. Fly.io Machines is not marketed at agent builders and has no SDK worth the name, but it gives you a real Firecracker VM with a durable volume, a public IPv4, a Sydney region, per-second billing and nothing charged while stopped — for roughly a tenth of what the agent-native vendors charge for equivalent compute. A weekend of glue code buys a lot. Vercel Sandbox and Cloudflare Sandbox are the same argument at the platform level: they exist so the sandbox is a line in the app you already deployed, not a second vendor, a second bill and a second on-call rotation. Neither is fastest or cheapest, and neither is trying to be.

Where this goes next is consolidation on two axes. Snapshot-and-fork is becoming table stakes, which kills the cold-start pitch as a standalone product — if everyone restores a paused VM in 200 ms, the differentiator moves to what the VM can reach on the network. Expect egress control, per-sandbox audit logs and outbound proxies to become the thing vendors argue about by 2027, because that is the feature enterprise buyers block deals on. The providers with a real answer there today — Fly, Northflank, Cloudflare, and anyone self-hosting Firecracker — are better positioned than their boot times suggest.

What cold start actually means

There are at least five distinct events a vendor might be timing, and the marketing word for all of them is "cold start".

The first is snapshot restore: a memory image of a booted VM mapped back in on a host that already has the pages. This is where the sub-100 ms numbers come from, and it is genuinely fast — but it presumes hardware allocated, image cached and pool warm. The second is guest boot: Firecracker to userspace, historically about 125 ms, which is where the whole microVM category gets its reputation. The third is scheduling: finding a host with capacity, which is zero when the pool is warm and seconds when it is not. The fourth is image pull, zero for a cached base image and thirty seconds for your 4 GB CUDA layer. The fifth is the API round trip and TLS handshake, 20 ms from the same region and 300 ms from Sydney.

Vendors quote the smallest of these. Agents pay all five. That is why the measured column here starts from a client, on an account with nothing running, and stops when a trivial command has returned output — and why it reads slower than every homepage. Any provider whose measured p50 lands within 3× of its claim is doing well, and most of this roster clears that bar: E2B at 2.0×, Modal at 1.5×, CodeSandbox at 1.2×, Fly at 3.0×, Daytona at 3.9×. Exactly one provider is an order of magnitude out — Blaxel, whose 25 ms is a snapshot restore and whose estimated create-to-output is 20× that. The size of a vendor's gap is a better guide to how it defines the word than to how fast its infrastructure is.

The pricing traps

Hourly compute is the number everyone compares and rarely the number that dominates the bill.

Watch the meter's units first. Per-second billing across 400 sandboxes that each live 12 seconds is a completely different bill from per-minute billing on the same workload, because per-minute rounds every one of them up. Cloudflare's 10 ms granularity and Vercel's active-CPU model are the friendliest shapes for bursty agent traffic; anything billing in minute blocks punishes exactly the pattern agents produce.

Then watch idle. An agent session is mostly the sandbox waiting for a model to finish generating. If the provider bills wall-clock, you are renting a CPU to wait on someone else's inference queue. Providers that bill only active CPU, or let you pause to storage-only rates, can be several times cheaper on identical work despite a higher headline rate — check the pause API's latency before relying on it, because pausing something you resume 400 ms later is a false economy.

Finally, the line items nobody puts on the pricing page: snapshot storage charged per GB-month and never garbage-collected, egress at cloud rates when your agent downloads a model, volumes that outlive the sandbox that created them, and platform minimums such as a paid Workers plan or a monthly SDK tier that make a "$0.05/hr" provider cost $150 before the first sandbox boots.

How to choose in an afternoon

Start with the isolation requirement, because it is the only one that cannot be retrofitted. Untrusted third-party code means a kernel boundary — microVM or gVisor — and it means outbound network control, which immediately shortlists Fly, Northflank, Cloudflare, E2B and self-hosted Firecracker. Your own agent's code on your own data is a much lower bar and every provider here clears it.

Then decide whether you are buying infrastructure or buying a product. If the sandbox is your product's core loop, take the agent-native SDK and the snapshot semantics — E2B, Daytona, Runloop — and treat the per-hour premium as the price of not maintaining an orchestrator. If the sandbox is one feature inside something bigger, use whatever your platform already gives you and revisit when the bill or the boot time starts to hurt.

Then run one benchmark before you sign anything: your real image, from your real region, 50 cold creates, recording p95 rather than p50. p50 tells you how the vendor's pool behaves on a good day. p95 tells you what your users see when the pool is cold, and that is the number deciding whether your agent feels instant or broken. If a vendor will not let you run that test on a trial account, that is itself the answer.

What is sandbox providers compared: cold start, isolation and price?

Sandbox providers compared: cold start, isolation and price

On toolweight, Sandbox providers compared: cold start, isolation and price means the 20 tools benchmarked on this page — E2B, Modal, Daytona, Fly.io Machines, Cloudflare Sandbox, Vercel Sandbox, CodeSandbox SDK, Northflank, Runloop, Blaxel, Freestyle, Val Town, Riza, Namespace, GitHub Codespaces, box, ascii, exe.dev, Islo, Self-hosted Firecracker — judged on the same 26 fields, from the same sources, on the same date. The question it exists to answer: Which sandbox provider should I run untrusted or agent-generated code on?

How does toolweight compare these?

Field values come from vendor documentation, pricing pages and public changelogs, each cell carrying its own provenance. Prices are the list rate for roughly 2 vCPU and 4 GB with no committed spend, normalised to an hourly figure. Several vendors bill per second or per credit and never publish an hourly number, so those cells are marked inferred with the arithmetic noted, or left null where a guess would be worse than saying nothing.

Cold start is the reason this page exists, so read the two columns together. Claimed cold start is whatever the vendor puts on its own homepage, and it is scored at weight zero on purpose — it is a marketing artefact, not a measurement. Vendors quote wildly different things under the same word: time to restore a memory snapshot on already-provisioned hardware, time for the guest kernel to boot, time for the API to return a sandbox ID, or p50 across a pool deliberately kept warm. Measured cold start is what toolweight publishes from its own harness: a single client in a fixed region, TLS handshake included, calling create-sandbox and blocking until a trivial command returns output — the full round trip an agent actually pays for. It runs against an account with no sandboxes alive, after a deliberate idle period, so nothing is pool-warm, and reports p50 and p95 over at least 50 runs so one good afternoon on a vendor's cluster cannot flatter the number. Until that harness has run against every provider here, cells in that column are estimates flagged as inferred, and no cell on this page currently carries measured confidence. We would rather show an honest estimate than launder a vendor's number into a benchmark.

Four rows come from a different kind of source and are flagged accordingly. The capability cells for box, ascii, exe.dev and Islo are transcribed from the comparison table box publishes on its own site, at box.ascii.dev/compare. That is the least neutral source on this page: a chart drawn by a vendor, with that vendor's two products in the first two columns and its competitors arranged around them. toolweight has verified none of it. Every field the table does not address — price per hour, cold start, isolation technology, persistence, regions, SDKs, funding, stars — is left null and unknown rather than filled in by inference.

Three rules keep that table from being laundered into evidence it is not. First, only a cell the table states directly is marked vendor-claimed; where a value had to be derived — reading root access off a "Docker inside the VM" row, or a preview-URL rating off two rows about IP addresses — the cell is marked inferred, because a sound derivation from a vendor's claim is still a derivation. Second, a marketing row is not a measurement: "runs 24/7" is not a published runtime ceiling and "1000+ concurrent VMs ergonomically" is not a published concurrency limit, so those cells are unknown rather than awarded this page's maximum. Third, and most important, absence from a row is scored asymmetrically on purpose. A vendor leaving its own product out of a row is a concession against interest and is credited — box and ascii are both absent from the process-fork and sub-500 ms rows, and that is recorded. A vendor leaving a rival out of a row is the weakest class of claim on this site, so it is recorded as nothing at all: exe.dev and Islo are missing from the snapshot rows, and their snapshot-and-fork cells stay unknown rather than becoming a scored zero on a competitor's say-so. These four rows stay in this shape until they can be re-sourced from each vendor's own documentation, at which point the provenance on every cell changes with them.

Full methodology and sourcing policy →

Frequently asked questions

Is a container enough to run LLM-generated code, or do I need a microVM?

If the code comes from your own model and your own prompt, a hardened container is usually fine. If it comes from a user, assume a container escape is a matter of time and budget: shared-kernel isolation has a steady drip of CVEs. A microVM (Firecracker, Cloud Hypervisor) or gVisor gives you a syscall boundary a kernel bug in the guest cannot cross. Check what you are actually buying rather than what the homepage implies — Daytona, the quickest-booting agent-native provider here, runs Docker containers on a shared kernel by default, and several vendors describe a container as having a "dedicated kernel" when they mean a dedicated namespace. The price difference is now small enough that the container answer is rarely worth defending.

Why is the cold start I measure so much slower than the number on the vendor's homepage?

Because they are measuring different things. Vendor numbers typically time a snapshot restore on hardware already allocated to you, excluding TLS setup, API scheduling, image pull and the round trip back to your client. Your number includes all of it, plus whatever queueing happens when their pool has scaled down. A 90 ms claim routinely lands between 400 ms and 2 s in production, and the gap widens the further you sit from their region.

What does snapshot and fork actually buy an agent?

Branching. You run a build, snapshot it, then fork the snapshot five times to try five patches in parallel without repeating the setup. It also makes idle cheap: pause mid-session, pay storage only, resume in a fraction of a cold boot with the process tree and page cache intact. For long agent sessions that is worth more than raw boot speed, because you only cold-start once.

Can I run a browser inside the sandbox instead of paying for a browser API?

Yes, and for a handful of pages a day it is cheaper. Chromium in a sandbox with a CDP endpoint costs compute you are already paying for. What you do not get is residential egress, CAPTCHA handling and the fingerprint work dedicated providers do, so scraping targets with real bot defences will fail. Use the sandbox for your own apps and internal tooling; use a browser API when the site is fighting back.

How locked in am I if I pick the wrong provider?

Less than it feels. Every provider here exposes the same four verbs — create, exec, read/write files, kill — so a thin interface of your own keeps swapping down to about a day of work. The real lock-in is elsewhere: proprietary image formats, snapshot artefacts that only restore on the vendor's cluster, and preview-URL domains baked into links your users have already saved. Keep images as plain Dockerfiles and you keep your exit.

Which providers can run in Sydney or another Australian region?

Fly.io Machines has a Sydney region and GitHub Codespaces offers an Australian region; Northflank reaches one through bring-your-own-cloud. The agent-native vendors — E2B, Daytona, Vercel Sandbox, Runloop — are US and EU only at the time of writing, which adds 150–250 ms of round trip to every API call made from Australia. If your control plane runs in Sydney, that latency dwarfs the cold-start differences everyone argues about.

Do I pay while a sandbox sits idle waiting for the model to respond?

It depends on the meter, and this is where bills go wrong. Fly charges nothing for a stopped machine beyond storage; E2B and CodeSandbox bill storage on a paused sandbox; Vercel bills active CPU, so a sandbox blocked on the network costs very little. Providers billing wall-clock per second charge you for every token the model is still generating. On agent workloads that idle time is usually the majority of the session.