Add jspi-hooks pass - #9102
Add jspi-hooks pass#9102guybedford wants to merge 1 commit into
Conversation
2e83607 to
8e90b89
Compare
Using separate functions seems like it would be better for follow-up optimizations and seems like it might be simpler in the end than using opaque operation codes. Stepping back, is the idea that this would be used for "advanced" JSPI usage like for fibers and multi-stack reentrancy? It seems that these use cases all boil down to some version of green threads. Will "plain" JSPI usage continue to not need this pass? Would it simplify things at all if we can assume that there is only a single, blessed JSPI suspension import? All other "suspending" imports could simply take the returned promise as an externref and then pass it to the single cc @sbc100. |
8e90b89 to
4577ec3
Compare
|
Thanks for the feedback @tlively I've updated the PR to use 4 separate hook functions instead.
Exactly, in particular for Cloudflare we need to support reentrant JSPI with multiple Tokio runtimes via JSPI-scoped thread locals in the same Wasm instance so that durable objects can serve multiple IO-segmented requests. So yes exactly supporting green thread primitives or otherwise via these hooks.
It might reduce the import interface at the cost of the double call and indirection. But imports are already a short list, and the problem is exports listing is the long one that would still be required so it wouldn't really fix the whole problem I don't think. |
Wraps the JSPI boundary of a module with calls to fiber lifecycle hooks that the module itself exports: promising exports (matched by the jspi-exports patterns) call __jspi_enter before and __jspi_exit after the wrapped function, and suspending imports (jspi-imports patterns) call __jspi_suspend before and __jspi_resume after the import. The hooks must run inside the fiber's own wasm frames at the instruction before/after the boundary call, which a JS wrapper cannot do as it only observes the transition a microtask later, when another fiber may already have run. The wrappers keep all state in locals: the i64 token returned by the "before" hook is forwarded to the "after" hook and is otherwise opaque to the pass (Emscripten uses a pointer to its fiber record, hence i64 so that memory64 pointers fit). Exceptional exits call the "after" hook with error=1, then rethrow unchanged (catch_all_ref/throw_ref), so the hook sees the fact of the error but the pass needs no knowledge of any tag. The legacy try/catch form is emitted when the module already uses legacy EH, since engines reject mixing it with try_table, and try_table/exnref otherwise; a module that already mixes both is rejected. Imports are wrapped by moving the import to a new function and turning the original function object into the wrapper, so all existing uses (calls, ref.func, element segments, exports) reach the wrapper directly. With --pass-arg=jspi-dyncalls the pass also exports a wrapped __jspi_dyncall_<sig> trampoline around call_indirect for every function signature in the table, so that hosts making function pointers promising (Emscripten's promising dynCall and embind async functions) do not bypass the hooks.
4577ec3 to
d6483a0
Compare
Are these expected to literally be |
|
No plans to literally remap all thread locals. I suspect we'd want to introduce a new Rust macro utility like |
Backport the JSPI lifecycle hooks, REENTRANT_JSPI fiber stacks and epoll listener API to the 6.0.9 frontend, and take Binaryen from a release carrying the jspi-hooks pass (WebAssembly/binaryen#9102). The release sysroot stamp is dropped after patching so emcc installs the new headers.
|
How does this relate to the co-operative threading ABI that WASI is proposing: https://github.com/WebAssembly/wasi-sdk/blob/main/CoopThreading.md. Could JSPI piggyback on this instead of creating a separate standard? |
Very interesting, I hadn't seen that, I did an AI guided deep-dive on the topic and got it to the following response which seems to match my own intuition in that Jco might benefit from the same hooks: _The coop-threading ABI already has the same shape on the export side: wit-component wraps every lifted export/callback/dtor with an in-wasm __wasm_task_hook(kind) call, and wasi-libc's hook allocates the task's stack and TLS and returns the SP to install — that's __jspi_enter/_jspi_exit. What it doesn't have is a suspend/resume hook around blocking calls, because the engine holds the SP/TLS in per-thread context slots that survive a block. JSPI has no such slots; the SP/TLS cells are wasm globals, so the switch has to happen in-fiber at the suspend/resume boundary with the identity carried in a wasm local. That's the whole delta: the import half plus the token (and an error flag, since JS/C++ exceptions cross the wrapper). So I don't see this as a separate standard — the hook events and the runtime policy (stack+TLS per task, TLS kept across suspension, freed at exit) are the same, and I'd want the kinds aligned so one libc implementation serves both. The codegen half of the coop ABI (libcall-thread-context) doesn't help under JSPI: its purpose is to let the engine own the cell, and here the global already is the cell. The reverse also holds: a component hosted on JSPI needs the import-half hook too. After a resume the first thing a p3 task does is context.get 0 for its stack pointer, and a JS-side "current task" register can't be updated soundly at resume time (two fibers' continuations can run before either resume job). jco already runs core modules through wasm-opt (its asyncify mode did exactly this), so it could apply the same pass — or wit-component's fixup could grow the import half natively, which is probably the right long-term home for the CM. Either way the hook contract is what's worth sharing; the instrumenter can differ per toolchain. Of course, having feedback from @vados-cosmonic would help here to figure this out further too. |
Backport the JSPI lifecycle hooks, REENTRANT_JSPI fiber stacks and epoll listener API to the 6.0.9 frontend, and take Binaryen from a release carrying the jspi-hooks pass (WebAssembly/binaryen#9102). The release sysroot stamp is dropped after patching so emcc installs the new headers.
|
I was able to convince myself that we need cleanup hooks on at least the export side of things. When Wasm returns from a JSPI'd export, the export promise is resolved with its return value. But then any cleanup logic chained onto that promise is pushed to the end of the microtask queue, so it could be preempted by another task already on the microtask queue. Such a task could call into Wasm and observe the not-yet-cleaned-up global state from the completed task. There's no other way to clean up after an export returns besides by chaining the cleanup logic onto its promise. If jco plans to lower the component model's It should still be possible to use Wasm wrapper functions around the original exports, but those would have to be generated offline, in which case they are no different than the hooks approach, or would have to be generated dynamically based on the types of the exports, which sounds unpleasant and would not scale to Wasm GC. |
|
Hi @tlively so our co-op threads implementation ( I'm not sure that we need this hooks pass to implement co-op threads on the Component Model side, because of the use of context builtins in particular... I'm not sure this is relevant there, but I'm still digesting this so I could be wrong. |
|
@vados-cosmonic the thinking here is that under JSPI, the Wasm stack and the shadow LLVM stack need to be kept in sync since JSPI will just switch the Wasm stack not the shadow LLVM stack. And the way to ensure they are correct is to always instrument all functions that are directly exposed in JSPI exports or imports. That instrumentation can either be on the bindings side for the system, or equivalently via something like these hooks. But because this stack alignment problem is unique to JS, I would be surprised if it just worked with reentrant JSPI without some extra machinery being needed, unless the component model bindings already do a full stack save and restore of the linear memory shadow stack around a JSPI suspension. |
|
Thanks for the explanation Guy, that makes it a bit easier to work out the concern here -- specifically dealing with the guest shadow stack, one thing missing from the discussion here is the work that TartanLlama & others already added to LLVM to enable upstream support for this: llvm/llvm-project@577e9a7 This is what I was referring to by the note about context.get/set, I had to look it up since I'm certainly not an expert in the wasi-libc implementation but the previously mutable single stack pointer becomes calls to For example in the new Rust target
IF I'm understanding right, as long as the version of LLVM is recent enough, we don't need the extra hooks here on the CM side (i.e. a recent-enough LLVM will be using the wasm builtins not However, in the case of Wasm w/ mutable stack pointers, IIRC the rewrite approach is what is currently done by There's one new option though I think isn't represented here, which is adding the relevant configuration & filling out rewriting in So AFAICT we do need something like
So essentially, under P3 yes, this is what happens and it's up to the guest toolchain. BUT for the the mutable stack pointer case, yeah it's not going to work out of the box. |
* worker-build --emscripten Build Workers for wasm32-unknown-emscripten through a worker-build provisioned toolchain: the pinned emsdk release is installed into the cache directory and the frontend patched with backports the Rust link needs (marker-based -sWASM_BINDGEN, NODERAWSOCKETS DNS). rustc drives emcc as the linker for a bin target, emcc runs wasm-bindgen post-link, and the output is collected into the standard build/index.js shape. The worker crate's start function becomes private, since its public init export collided with Emscripten's own. * Keep DWARF in emscripten builds The exnref translation failure on debug builds was wasm-bindgen 0.2.128's --keep-debug output (walrus 0.27.1 .debug_loc), fixed on main by #5328, not binaryen on the linked module. The CLI floor already covers it. * Skip reinit on emscripten and support minified emcc output The worker crate no longer registers the abort reinit hook on wasm32-unknown-emscripten, so the released wasm-bindgen CLI builds --emscripten --release; only debuginfo builds still need #5328. Export and wasm import detection now handle emcc's minified release output. * worker-build: leave every cloudflare: module external when bundling The runtime serves the whole cloudflare: namespace (cloudflare:node, cloudflare:test, ...), not only the three that were listed, and a wasm-bindgen module import of one of them must reach the output as an import rather than fail to resolve in esbuild. * Reentrant JSPI toolchain Backport the JSPI lifecycle hooks, REENTRANT_JSPI fiber stacks and epoll listener API to the 6.0.9 frontend, and take Binaryen from a release carrying the jspi-hooks pass (WebAssembly/binaryen#9102). The release sysroot stamp is dropped after patching so emcc installs the new headers. * emscripten-tcp: note reentrant activations * Use wasm-bindgen JSPI lifecycle hooks integration wasm-bindgen/wasm-bindgen#5333 instruments jspi exports and suspending imports with emscripten's fiber hooks, so under -sREENTRANT_JSPI each fetch activation runs on its own stack. The example takes wasm-bindgen from the submodule until it is released. * ci: build the emscripten example with the submodule wasm-bindgen * ci: accept any HTTP/1.x status line in the emscripten smoke test * Fiber-owned Tokio context under REENTRANT_JSPI Pass --cfg tokio_jspi_hooks so Tokio's emscripten port registers the JSPI lifecycle hooks, and move the example to that Tokio branch. Concurrent fetch activations each drive their own runtime. * worker-build --emscripten: derive exported classes from DurableObject The runtime exposes RPC on a Durable Object class only when it derives from cloudflare:workers' DurableObject; the wasm-bindgen classes are plain, so splice its prototype in. * JSPI exports for event and Durable Object handlers on emscripten On wasm32-unknown-emscripten the handler wrappers are #[wasm_bindgen(jspi)] exports suspending on the handler future, so each activation is its own fiber. The wrappers are left for rustc to expand: expanding both cfg variants in the proc macro deduplicated their export descriptors. * worker-build --tokio: event-loop and JSPI Tokio integration `--emscripten --tokio` schedules each handler on a Tokio event-loop runtime per invocation (tokio-rs/tokio#8479), whose wait is the host event loop, so plain async handlers use tokio::net and tokio::time with no stack switching. `--emscripten --tokio=jspi` is the previous behaviour: JSPI exports on their own fibers blocking on a runtime. The mode reaches the macros as cfg(worker_tokio = ...). The epoll listener backport now delivers readiness through a macrotask: workerd drains microtasks synchronously inside builtin module loads, so a microtask delivery ran one runtime's drive under another's connect(). * Update Cargo.lock for wasm-bindgen-futures tokio feature * Regenerate emscripten patches from the current upstream PR heads reentrant-jspi.patch now carries #27698, #27699 and #27547 at their heads (listener keepalive holds, teardown wakes, macrotask delivery) and noderawsockets-dns.patch carries #27693 at its head (node:dns only, no virtual /etc/hosts), both as exact diffs against the 6.0.9 tree. * Regenerate emscripten patches from the rebased upstream stack epoll-listeners.patch carries #27547 at 34e2cc057 and #27720 (timeout keepalive release), noderawsockets-dns.patch carries #27693 as rebased onto that, and the JSPI hooks / REENTRANT_JSPI backport (#27698, #27699) moves to a trailing jspi-hooks.patch that only --tokio=jspi depends on. * Track epoll listener PR heads: MINIMAL_RUNTIME guard around maybeExit * Record the rebaselined epoll listener PR heads * Track the epoll listener PR heads: native_sigs.py ordering * Drop the JSPI Tokio mode: --tokio is the host-driven event loop The event-loop integration never needed stack switching, so --tokio=jspi, its lifecycle-hooks backport (jspi-hooks.patch), the -sJSPI link, the Binaryen release carrying the jspi-hooks pass, the tokio_jspi_hooks cfg and the JSPI export variant in worker-macros go. --tokio is a plain flag and the emsdk release supplies the whole backend. The noderawsockets-dns patch stays: without stack switching getaddrinfo returns EAI_AGAIN for hostnames, a clean io::Error instead of the host's uncaught "Invalid IP address" from net.connect. * Move --tokio onto tokio#8484 host-driven event loops tokio-rs/tokio#8484 replaces EventLoopRuntime with EventLoop / LocalEventLoop; the example takes tokio from its branch and mio from its emscripten branch (tokio-rs/mio#1969). The wasm-bindgen submodule moves to the rebuilt emscripten stack: main plus the reinit fix (#5332), legalized export names and the tokio attribute (#5334) adapted to LocalEventLoop, without the emscripten JSPI support. worker's schedule_isolated use is unchanged. * Restore the wasm-bindgen path patches for the submodule stack 0.8.6 dropped them to release on stable wasm-bindgen; worker's tokio feature on wasm-bindgen-futures and #[wasm_bindgen(tokio)] come from the submodule until they ship. * Export the connect handler through async_export_mod Brings #[event(connect)] (#1041) in line with the other handlers so it runs on the Tokio event loop under --tokio. * Track the simplified epoll listeners and tokio#8484 runtime keepalive emscripten#27547 now delivers through unref'd handles with no listener keepalive of its own, and #27720 / #27693 are regenerated from their merged main commits. The example's tokio moves to 5ef5ab3d, where the event loop holds the runtime keepalive itself. * Move onto tokio#8484 hosted event loops The example's tokio moves to 067f92b6, where the emscripten event loop is built hosted: it schedules its own drives and holds the runtime keepalive while it has tasks. The wasm-bindgen submodule follows with the tokio attribute built on it; worker's use of schedule_isolated is unchanged. * Track the emscripten socket stack heads Regenerates the 6.0.9 backports from the current pull request heads: the epoll listeners are unchanged in content, the DNS patch gains emscripten_dns_lookup_async / emscripten_dns_lookup_result (#27742), and a new accept-blocking.patch carries the part of #27724 that reaches the single-threaded Node build, recv surfacing a pending socket error ahead of EOF. * Use #[wasm_bindgen(tokio = "isolated")] for event-loop exports Handlers are plain async exports; worker-build --tokio only adds the attribute option, and wasm-bindgen's own glue bridges the isolated event loop to the returned Promise. Removes worker::__tokio_promise and the Durable Object static_self escape, since wasm-bindgen holds the borrow across an async method. * Assert handler futures unwind-safe for panic = "unwind" async exports * Make the worker `tokio` feature the event-loop switch Replaces `worker-build --tokio` and the injected `worker_tokio` cfg: the `tokio` feature on `worker` enables `worker-macros/tokio`, which emits `tokio = "isolated"` on handler exports, and `wasm-bindgen-futures/tokio`. worker-build passes `--cfg tokio_unstable` for every emscripten build. * Split the emscripten examples: std::fs, Tokio, then TCP emscripten is a word-frequency counter over files on the in-memory filesystem with no Tokio; emscripten-tokio shows spawn, timers, timeout, mpsc, Mutex and join! in plain handlers under the tokio feature; emscripten-tcp becomes a small TCP proxy (POST body or HEAD request to an address, from the Worker or a Durable Object). CI builds and smokes all three. * Move to Emscripten 6.0.10 and the tagged Tokio stack worker-build installs emsdk 6.0.10, which ships the wasm-bindgen marker link, timeout keepalive release and NODERAWSOCKETS DNS; the remaining patches are #27547 and #27742 as carried by guybedford/emscripten 6.0.10-cf.emscripten. The examples pin the guybedford tokio 1.53.1-cf.emscripten and mio 1.2.3-cf.emscripten tags, and the plain emscripten example documents that without Tokio none of the patches are needed. wasm-bindgen renamed the export attribute to experimental_tokio and gates it behind --cfg wasm_bindgen_unstable_tokio; worker-build passes that and tokio_unstable only when the worker tokio feature resolves, so a Worker without it builds against stock Tokio. The new examples are also excluded from the workspace, which the previous commit missed. * Point wasm-bindgen at gbedford/emscripten-stack 2988e0efe wasm-bindgen main plus #5334 with the tagged Tokio and mio pins. The Examples job skips the three emscripten examples, which the dedicated job builds. * Follow the sealed wasm-bindgen and Tokio tips wasm-bindgen moves to gbedford/emscripten-stack ea4368ac0 (main plus #5333) and the examples' lock to the 1.53.1-cf.emscripten tag's current commit. Neither changes behaviour here: the JSPI layers are gated behind cfgs worker-build does not pass. * Point wasm-bindgen at main and pass the crate paths on the impl Main carries #5333 as merged, at the released 0.2.128 version, so the workspace and the CLI worker-build downloads stay aligned until 0.2.129 ships. With #5333's fix the Durable Object impl attribute can name the wasm_bindgen_futures and js_sys paths directly instead of importing them into scope. * udpate readmes * fixup * Resolve hostnames in the TCP example The Tokio patchset tag now routes ToSocketAddrs through Emscripten's asynchronous getaddrinfo (#27742, in the worker-build patches), so the example is the plain TcpStream::connect((host, 80)) HEAD probe again and takes hostnames from the Worker and the Durable Object; CI smokes both. * Bring the emscripten examples into the workspace as experimental_tokio The three examples are ordinary workspace members: the Emscripten patchset (Tokio, mio, libc, wasm-streams) joins the root patch block alongside the wasm-bindgen paths, and the examples keep only their dependencies. The worker feature is experimental_tokio, matching the wasm-bindgen attribute it emits, and is inert off emscripten so the workspace builds and lints for the host. Each README opens with the clone-and-run steps. * Keep the emscripten examples standalone with the Tokio patchset in place Each example's Cargo.toml carries the required Tokio, mio and libc patches as the reference for a Worker's own manifest; the wasm-bindgen path patches are the workspace's alone, so the examples take the crates from crates.io. The worker feature is experimental_tokio. * Move to wasm-bindgen 0.2.129 The workspace and submodule follow the release; the examples resolve it from crates.io and worker-build downloads its CLI, whose --keep-debug output debuginfo builds need. * Durable Object connect handlers and the emscripten-incoming-tcp example Stub::connect opens a TCP connection to an object, served by the new DurableObject::connect handler that #[durable_object(connect)] exports; Socket::handle_as_node_connection hands an inbound socket to the Node-style server listening on its port. Under experimental_tokio the object's handlers share the thread's ambient event loop, a Durable Object being one I/O context, while event handlers keep a runtime per invocation. The example runs a Tokio TcpListener line-echo server inside an object, with the Worker's connect handler piping inbound connections to it. * Regenerate the epoll listeners patch from the moved emscripten tag emscripten#27547 now schedules a host turn per wake instead of coalescing onto a pending immediate, which a host may drop: Workers cancels the immediates left over when a request settles, and the listener stayed silent thereafter. Inbound connections after the first now reach the object's listener. * Expose Storage::sync and the raw state and storage objects storage.sync() waits for issued writes to commit; State::as_raw and Storage::as_raw hand the underlying JavaScript objects to libraries that take them directly, such as a filesystem mounted on the object's storage. * Serve inbound TCP with a per-connection runtime off a shared listener The final Emscripten and Tokio tags deliver epoll readiness and hosted drives as microtasks, so a drive runs in the request whose event woke it. The incoming-tcp example is then a plain Worker: the connect export runs on the shared event loop, which hosts only the TcpListener and its accept loop, and each accepted connection is moved onto an event loop of its own inside the connect invocation that delivered it. The epoll listeners patch follows the tag. * Reject failed connect handlers, align example compat dates, take the accept-queue emscripten tag A connect handler returning Err now rejects the export promise instead of panicking, closing only that connection. The test for it is skipped until the bundled workerd carries cloudflare/workerd#7518. Examples move to compatibility_date 2026-09-02 with new_module_registry; emscripten-tcp drops its Durable Object probe. The epoll-listeners patch picks up the NODERAWSOCKETS accept queue of depth one from the tag. CI waits on wrangler's ready line rather than a fetch for connect-only Workers, and carries a disabled cold-burst check for the pending tokio sync-drive fix. * readme updates * Document that a connect handler's return closes the socket * worker-build: single Binaryen core on macOS Binaryen's worker threads overflow their stacks on large modules there.
This implements a
--jspi-hookspass that wraps the JSPI entry/exit/suspend/resume boundaries of a module with calls to corresponding lifecycle hooks the module itself provides.This is needed in Binaryen itself because it cannot be done at another layer of the system - suspension incurs a microtask, so that any JS wrappers around the suspend JS function will not be correct, and any exit wrapper on the WebAssembly.promising itself will also not have the correct timing.
Only by directly integrating the synchronous timing of the hooks into the Wasm can we get sound tracking of JSPI context switching.
The use cases are various here:
The alternative to this approach would be to more heavily define JSPI instrumentation:
Given where JSPI is in its ecosystem adoption as an emerging convention, a general hooks-based approach feels like the best way to start on these problems for now.
The hooks are four function exports which must be provided by the binary for the transform to work:
They are called on every JSPI enter/exit/suspend/resume operation:
Then similar to the lines of asyncify, the PR here implements for Binaryen that contract as a transform pass:
--pass-arg=jspi-exports@<patterns>: each matching export is retargeted to a wrapper calling__jspi_enterand__jspi_exitaround it.--pass-arg=jspi-imports@<patterns>: each matching import is moved to a new import and the original function object becomes a wrapper calling__jspi_suspendand__jspi_resumeand passing the token between them, so every existing use (calls,ref.func, element segments, exports) reaches the wrapper with no reference rewriting.--pass-arg=jspi-dyncalls(plusjspi-dyncall-sigs@ii,vifor extra signatures) also exports a wrapped__jspi_dyncall_<sig>(fptr, ...)trampoline aroundcall_indirectper function signature in the table, so hosts that make function pointers promising (Emscripten's promisingdynCall, embindasync()) do not bypass the hooks.<sig>is thegetSig()alphabet.When instrumented:
try_table (catch_all_ref)calls the after hooks witherror=1, thenthrow_refs the exception unchanged, whatever its tag, so the pass needs no knowledge of any tag and no imports. Modules already using legacy EH get the legacytry/catch_all/rethrowform, since engines reject mixing the two; a module that already mixes them is rejected with a pointer to--translate-to-exnref.*, comma/newline separators and@fileresponse files like asyncify.The actual behaviors of the JSPI system are otherwise entirely left to the hook runtime implementations, with Emscripten expected to support a C runtime implementation for interfacing with JSPI hooks and also using it to build
REENTRANT_JSPI(PR which is based on this).Due to the generality though, other runtimes could implement their own custom systems similarly.
Tests: lit coverage of import wrapping through direct calls,
ref.func, element segments and exported imports; export retargeting with internal callers untouched; hook exclusion; features; response files; both EH forms; multivalue; trampolines with explicit signatures; the Fatal cases;--roundtrip/-O2stability; and an engine-leveltest/lit/d8run of the wrapped module under real JSPI (V8) checking the full event trace, exception identity through the rethrow,SuspendErrorreaching the resume hook, id forwarding across a real suspension, nested promising entry from a plain import, and a trampoline call.Made with AI assistance under my review