Skip to content

fix(orchestrator): add opt-in template build concurrency cap - #3127

Open
renyuanc wants to merge 1 commit into
e2b-dev:mainfrom
renyuanc:renyuan/template-build-concurrency-cap
Open

renyuanc wants to merge 1 commit into
e2b-dev:mainfrom
renyuanc:renyuan/template-build-concurrency-cap

Conversation

@renyuanc

Copy link
Copy Markdown

Summary

TemplateCreate launches one detached goroutine per request with no concurrency limit. Each goroutine drives a full template build, extracting an ext4 rootfs and booting a provisioning Firecracker VM, so N simultaneous TemplateCreate calls put N rootfs assemblies and N micro-VMs on a single template-manager host at once, with no backpressure or queue. Under a burst (CI fan-out, batch builds, several teams at once) this oversubscribes host disk, I/O, and RAM.

This PR bounds concurrent builds per host with a weighted semaphore, gated by a feature flag.

Fixes #3070.

What this does

  • Adds MaxConcurrentTemplateBuilds (max-concurrent-template-builds, default -1) to flags.go, alongside the existing MaxConcurrent* family. This closes the one heavy concurrent path that lacked a cap (template builds).
  • Acquires the semaphore inside the detached build goroutine The handler must stay non-blocking so the API's synchronous gRPC trigger returns immediately while blocking it would surface to the caller as a request timeout. Excess builds queue (they already report Building, which the API status poll tolerates) instead of over-committing the host.
  • Kill switch: a non-positive value (<= 0, e.g. -1) disables the cap entirely (unbounded builds), matching the repo's -1 = unlimited convention (TCPFirewallMaxConnectionsPerSandbox, SandboxMaxIncomingConnections).

Testing

  • Built + go vet clean (linux/amd64, CGO).
  • Validated on a 2-node staging cluster: with the cap set to 2 and 10 builds fired simultaneously, concurrent builds held at exactly 2 per node (the rest queued as Building); with the kill switch (-1) a node ran 3 concurrently, confirming the cap is bypassed.

Scope / out of scope

This caps build concurrency. It deliberately does not change build lifecycle/timeout handling. One consequence worth noting for reviewers: queuing makes a build sit in Building longer, which interacts with the API's build-status timeout and the build-cache TTL - a build can be marked failed for queuing too long That's a pre-existing lifecycle gap that this cap amplifies; it should be tracked separately

@cla-bot

cla-bot Bot commented Jun 28, 2026

Copy link
Copy Markdown

We require contributors to sign our Contributor License Agreement, and we don't have @renyuanc on file. You can sign our CLA at https://e2b.dev/docs/cla . Once you've signed, post a comment here that says '@cla-bot check'

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The AdjustableSemaphore.Acquire implementation has a critical race condition that can lead to a permanent deadlock when the context is cancelled. Because context.AfterFunc runs its callback asynchronously, s.cond.Broadcast can be called before the acquiring goroutine actually enters s.cond.Wait. Since sync.Cond broadcasts are not buffered, the signal is lost and the goroutine blocks indefinitely. To fix this, the AfterFunc callback in resizable_semaphore.go must acquire the semaphore's mutex before calling s.cond.Broadcast to ensure the broadcast is synchronized with the wait state.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

// Queue builds inside the background goroutine so the gRPC handler stays
// non-blocking and the node does not oversubscribe build resources.
if s.buildLimitEnabled.Load() {
if err := s.buildLimiter.Acquire(ctx, 1); err != nil {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The AdjustableSemaphore.Acquire implementation has a critical race condition that can lead to a permanent deadlock when the context is cancelled. Because context.AfterFunc runs its callback asynchronously, s.cond.Broadcast can be called before the acquiring goroutine actually enters s.cond.Wait. Since sync.Cond broadcasts are not buffered, the signal is lost and the goroutine blocks indefinitely. To fix this, the AfterFunc callback in resizable_semaphore.go must acquire the semaphore's mutex before calling s.cond.Broadcast to ensure the broadcast is synchronized with the wait state.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is pre-existing in shared code, not introduced by this PR. Suggest to fix it in a separate PR.

@renyuanc

renyuanc commented Jul 1, 2026

Copy link
Copy Markdown
Author

@cla-bot check

@cla-bot cla-bot Bot added the cla-signed label Jul 1, 2026
@cla-bot

cla-bot Bot commented Jul 1, 2026

Copy link
Copy Markdown

The cla-bot has been summoned, and re-checked this pull request!

@ValentaTomas
ValentaTomas force-pushed the main branch 2 times, most recently from 5aad415 to d71980e Compare July 25, 2026 22:53

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Template builds are unbounded: TemplateCreate spawns one goroutine (and one FC VM + rootfs assembly) per request with no concurrency limit

1 participant