news_article.exe
📰

2026年最佳Agent沙箱:E2B、Daytona、Modal、Cloudflare与Vercel的冷启动、按秒计费及网络策略对比

Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel

2026年8月27日1 次浏览来源:MarkTechPost 阅读原文

Every agent that writes code needs somewhere to run it. That somewhere is now a product category with at least a dozen vendors, four incompatible billing models, and marketing pages that quote cold starts measured under conditions nobody publishes. This comparison fixes the units. It covers the five platforms most teams shortlist — E2B, Daytona, Modal Sandboxes, Cloudflare Sandbox SDK, and Vercel Sandbox — along with Runloop, Fly.io Sprites, and Northflank where they change the answer. The four questions that actually decide this Feature matrices for this category are mostly noise. Four properties change architecture, and everything else is a preference: Cold start under concurrency: An agent loop that creates a sandbox per tool call pays this tax thousands of times a day. Filesystem...

Best Agent Sandboxes in 2026: Cold Start, Per-Second Pricing, and Network Policy Across E2B, Daytona, Modal, Cloudflare, and Vercel

Every agent that writes code needs somewhere to run it. That somewhere is now a product category with at least a dozen vendors, four incompatible billing models, and marketing pages that quote cold starts measured under conditions nobody publishes.

This comparison fixes the units. It covers the five platforms most teams shortlist — E2B, Daytona, Modal Sandboxes, Cloudflare Sandbox SDK, and Vercel Sandbox — along with Runloop, Fly.io Sprites, and Northflank where they change the answer.

The four questions that actually decide this

Feature matrices for this category are mostly noise. Four properties change architecture, and everything else is a preference:

Cold start under concurrency: An agent loop that creates a sandbox per tool call pays this tax thousands of times a day.

Filesystem persistence between turns: Does turn 2 see the pip install from turn 1, or does the agent rebuild its world?

Egress policy: Can the sandbox reach the internet, can you turn that off, and can you change your mind mid-session?

Idle billing: Agents spend most of their wall-clock waiting on a model. Somebody is paying for those seconds.

1. Cold start: what the numbers actually say

The vendor claims are not comparable to each other. Daytonas pricing page advertises sub-90ms sandbox creation. E2B is commonly cited at roughly 150ms. Modal advertises sub-second cold starts for pre-cached containers. None of these state concurrency, region, image size, or whether the clock stops at API acknowledgment or at first executed command.

The most useful public dataset is ComputeSDKs sandbox leaderboard, which is open source and runs on a schedule. It measures Time to Interactive (TTI): elapsed time from create() to the first successful command inside the sandbox, 100 iterations per provider, launched concurrently in a single burst, from a 4 vCPU host in Northern Virginia.

Results from the August 21, 2026 run:

ProviderMedian TTIP95P99Success rateVercel Sandbox0.67s1.04s1.12s100%Modal0.88s1.00s1.08s100%Runloop0.89s3.27s3.50s100%E2B1.61s1.77s1.81s100%Cloudflare5.06s6.04s6.48s100%Daytona0.27s0.43s0.44s37%

Three things in that table matter more than the ranking.

Burst is not the same test as sequential: Daytonas fastest published median is real, and on an earlier provider-page run it created sandboxes at a 0.10s median when launched one at a time. On the August burst run it posted the fastest median in the field and completed 37 of 100 attempts. A median you only reach on a third of your calls is not a latency number, it is a capacity number. Retry logic is not optional on any of these platforms.

Tail latency is the number to design against: Runloops median and Modals median are 10ms apart. Runloops P95 is 3.3x Modals. If your agents UX budget is one second, the median tells you almost nothing.

Cloudflare is measuring a different product: Sandbox SDK sits on Cloudflare Containers, which schedules a container instance and boots an image. That is architecturally a heavier operation than resuming a pre-warmed Firecracker VM, and 5s medians reflect it. Cloudflares own GA post is candid about the shape of the problem: booting a sandbox, cloning a repo, and running npm install takes about 30 seconds, while restoring the same environment from a backup takes about two.

The task worth measuring is the one your agent runs, not echo hello. A useful harness runs the same unit of work everywhere: install pandas, read a CSV, plot it, return a PNG. Time four checkpoints separately.

Copy CodeCopiedUse a different Browser

Report tti and task separately. Vendors optimize the first and readers care about the second. Pin the region, pin the image, and publish both the sequential and the concurrent series, because they answer different questions.

2. Per-second pricing, normalized

Published rates as of August 27, 2026, converted to a common unit. Modal prices per physical core, which it defines as 2 vCPU, so the vCPU-equivalent is shown for comparison.

PlatformCPUMemoryBilling basisPlan floorE2B$0.0504 / vCPU-hr$0.0162 / GiB-hrWall-clock, per secondFree Hobby; $150/mo ProDaytona$0.0504 / vCPU-hr$0.0162 / GiB-hrWall-clock, per secondNone; $200 creditModal Sandbox$0.1419 / core-hr (~$0.0710 / vCPU-hr)$0.0240 / GiB-hrmax(request, actual), per secondFree Starter; $250/mo TeamVercel Sandbox$0.128 / vCPU-hr active CPU only$0.0212 / GB-hr provisionedSplit: CPU active, memory wall-clockHobby allotment; Pro creditCloudflare Sandbox$0.072 / vCPU-hr active CPU only$0.009 / GiB-hr provisionedActive CPU + provisioned memory/disk$5/mo Workers PaidFly.io Sprites$0.07 / CPU-hr$0.04375 / GB-hrActive use only; sleeps when idleSubscription tiersRunloop$0.108 / CPU-hr$0.0252 / GB-hrRunning state; suspended is storage-onlyFree Basic; $250/mo ProNorthflank$0.01667 / vCPU-hr$0.00833 / GB-hrAllocated resources, per secondFree Sandbox tier

Two footnotes that people get wrong.

Modals sandbox tier is roughly 3x its standard Function rate ($0.00003942 vs $0.0000131 per core-second), and region selection adds 1.5–1.75x on top. Sandbox pricing is not Modals headline compute pricing.

Daytonas GPU rates are widely reproduced at $3.95/hr for an H100. Its live pricing page lists on-demand H100 at $2.27/hr and H200 at $2.61/hr. Third-party comparison tables in this category go stale within a quarter.

Rates are not costs. The model below fixes the workload and runs it through each rate card.

Assumptions: 2 vCPU / 4 GiB sandbox, 1,000 executions, no plan floor included, no egress, default region (Vercel iad1, Cloudflare standard-3 at 2 vCPU / 8 GiB / 16 GB disk since instance sizes are fixed).

Scenario A: short burst — 90s alive, 50% average CPU

PlatformCost / 1,000CompositionNorthflank$1.67$0.83 CPU + $0.83 memoryCloudflare$3.70$1.80 CPU + $1.80 memory + $0.10 diskE2B / Daytona$4.14$2.52 CPU + $1.62 memoryVercel$5.32$3.20 active CPU + $2.12 memoryModal$5.95$3.55 CPU + $2.40 memoryFly Sprites$7.88$3.50 CPU + $4.38 memoryRunloop$7.92$5.40 CPU + $2.52 memory

Scenario B: idle-heavy — 10 min alive, 5% average CPU

This is what a real agent loop looks like. The sandbox is open, the model is thinking, nothing is running.

PlatformCost / 1,000Change vs ANorthflank$11.116.7xCloudflare$13.873.7xVercel$16.273.1xE2B / Daytona$27.606.7xModal$39.666.7xFly Sprites (kept awake)$52.506.7xRunloop (kept running)$52.806.7x

Vercel moves from 4th-cheapest to 3rd, and its CPU line drops from $3.20 to $2.13 while everyone elses scales linearly. Cloudflares active-CPU line falls to $1.20. That is the entire argument for active-CPU billing, and it is worth roughly 2x on this workload.

The platforms that lose Scenario B can win it back, if your orchestration suspends between turns instead of holding the box open. Same workload, 30s awake per execution:

PlatformCost / 1,000MechanismE2B (auto-pause)~$2.16Pause costs ~4s per GiB of RAM, resume ~1s (docs)Fly Sprites$2.62Idle monitor sleeps the sprite within secondsRunloop$2.64Suspend stops compute billing; storage continues

E2Bs number includes ~17s of pause and resume overhead for a 4 GiB sandbox. That overhead is the deciding variable: pausing is only economical when the gap between turns is meaningfully longer than the pause itself.

Flys idle detector is specific about what counts as activity: an in-flight HTTP or API request, output to a sessions stdout, an open TCP connection, or an active task (sprites.dev). An agent that holds a connection open while it waits is an agent that is billed. Redirecting output to a file does not count, which is a real lever.

4. Filesystem persistence between turns

This is where the platforms diverge most, and where the wrong choice shows up as a rebuilt node_modules on every turn.

PlatformDefault on stop/idleMemory stateMechanismE2BonTimeout defaults to killPause preserves RAM and running processespause() / connect(), paused

> 分享: