Repository navigation
Native D3D12: build pipelines at first use, in parallel, and share id… - #139
Merged
Merged
Conversation
…entical libraries Branch: `pr/d3d12-pipeline-build` (one commit). Companion PR: `pr/d3d12-persistent-shader-cache`. ## Problem Ghost of Tsushima (Nixxes port) creates every pipeline of a level up front: about 14,000 at New Game, with 29,297 shader libraries. On an iPhone 17 Pro Max the loading screen spun until iOS killed the app. The last footprint line read 7.8 GB, with Metal's `currentAllocatedSize` at 5.1 GB, of which only ~620 MB were textures and buffers. Once that was fixed by building pipelines at first use, the next wall was the start of gameplay. `ExecuteCommandLists` went from 25 ms to 273 ms to 2.3 s per frame with the GPU 2-7 % busy (Metal HUD: 100 % of frame time encoding), and the game's watchdog stopped it. ## Cause - `CreateGraphicsPipelineState` / `CreateComputePipelineState` call `newRenderPipelineState` / `newComputePipelineState` immediately. Metal compiles each pipeline to GPU code and allocates for it at creation, even though a scene draws only a small fraction of them. - `mad_convert_stage_opts` creates a new `MTLLibrary` + `MTLFunction` for every stage of every pipeline. The game's 30,370 libraries came from only ~11,400 distinct converter outputs, and Metal keeps GPU-side storage per library. - Built lazily, pipelines compile on the submitting thread, one after another, at the first draw. The first gameplay frames need hundreds. ## Change All in `madeira-d3d12/src/pe/madeira_d3d12.c`: - **Lazy pipelines.** A plain render pipeline (no GS, no hull/domain) keeps its `WMTRenderPipelineInfo` (and vertex descriptor) and is built by `mad_pso_realize` at its first draw. A compute pipeline is built by `mad_cpso_realize` at its first dispatch. Each pipeline has its own `SRWLOCK` (zeroed by `calloc`), so builds of different pipelines run concurrently. If Metal rejects a pipeline at that point, the failure is logged (first 8) and its draws/dispatches are skipped, like the existing placeholder pipelines, instead of `CreateGraphicsPipelineState` returning `E_FAIL`. GS and tessellation pipelines are unchanged (eager). - **Parallel first use.** At the start of `mad_ecl_run`, `mad_prebuild_lists` collects the not-yet-built pipelines that the batch's lists bind (up to 256, skipping lists the replay will reject). When there are at least 2, it builds them on the calling thread plus a persistent pool of 3 worker threads (256 KB stacks, created once). Creating Wine threads per batch cost an 8 MB stack, a TEB and an emulator thread state each time. Batches use the pool one at a time. - **Shared Metal libraries.** `mad_convert_stage_opts` hashes the converted metallib bytes plus the entry name (128-bit key). If an identical library exists, it hands out the existing `MTLLibrary`/`MTLFunction` with a retain of each, so `pso_Release` and `mad_tess_free` stay balanced. The open-addressing table keeps one reference per distinct library, grows at 70 % load and never fills. - Switches (madeira.cfg, both default 1): `pso-lazy = 0` restores eager creation, and `pso-parallel = 0` turns the pre-replay parallel build off. Both are kept because they change when pipeline creation (and its failures) happen, which is useful for A/B runs. `ConfigCatalog.generated.swift` gains the two entries exactly as `build/tools/gen-config-catalog.py` emits them. Only those two lines are added, because the generated file on `main` is already stale for unrelated keys. - Log lines, in the file's existing style: - `pipelines are built ... (madeira.cfg pso-lazy)` once - `lazy pipelines built at first draw: N` every 500 - `pso-parallel: built N pipelines on T threads before replay` (first 8 batches, then batches of 32 or more) - `shared shader libraries: N reuses, M distinct` every 2000 ## Evidence From the fork, on an iPhone 17 Pro Max (iOS 27) with Ghost of Tsushima: - With lazy pipelines plus shared libraries (and the fork's persistent shader cache), New Game no longer runs out of memory. About 500 pipelines were built at first draw in the first scene, and the game reaches gameplay (~20-40 FPS at that point). - Parallel prebuild: the menu dropped to ~2 ms of `ExecuteCommandLists` per frame (45-48 FPS), with `pso-parallel: built 20..69 pipelines on 4 threads`. The per-batch thread version caused a storm of thread setups and failed 8 MB reserves at gameplay start. The persistent pool removed it. - Remaining cost: a pipeline used for the first time in a new area still causes a 50-550 ms hitch while Metal compiles it (no `MTLBinaryArchive` yet). ## Notes / risks - **The prebuilt `app/Madeira/arm64ec-windows/madeira_d3d12.dll` (and `d3d12.dll`) is not rebuilt in this PR.** It needs a rebuild with `build/madeira-d3d12/build-pe.sh`. I could not run llvm-mingw here. The file was syntax-checked with clang against mingw-w64 headers: no new errors or warnings compared with `main`, and a canary edit confirmed the new code is checked. - The shared-library table keeps every distinct library alive for the life of the process, even after all pipelines using it are released. That was the right trade for this game (11,400 shared vs 30,370 separate), but a title that streams many unique pipelines and frees them would hold them longer than before. - A pipeline that Metal rejects is now discovered at first draw rather than at creation, so the app gets a pipeline object whose draws are skipped instead of `E_FAIL`. `pso-lazy = 0` restores the old contract. - The lazy build reads `p->rps`/`p->cps` outside the lock on the fast path. They are written once, under the per-pipeline lock. - **Not ported:** - The fork's per-phase timing counters (`pso time ...`), which are diagnostics only. - The ntdll "wider high-band fallback" that shared a fork commit with the thread pool (`build/ntdll-unix/virtual_ios.c`), which is a separate memory concern. - The placeholder for hull+domain pipelines that cannot be built, which belongs with tessellation. - No submodule changes. --- 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0189oLHghpaYKLk4f786a6bc Signed-off-by: bahacan16 <190844990+bahacan16@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
…entical libraries
Branch:
pr/d3d12-pipeline-build(one commit). Companion PR:pr/d3d12-persistent-shader-cache.Problem
Ghost of Tsushima (Nixxes port) creates every pipeline of a level up front: about 14,000 at New Game, with 29,297 shader libraries. On an iPhone 17 Pro Max the loading screen spun until iOS killed the app. The last footprint line read 7.8 GB, with Metal's
currentAllocatedSizeat 5.1 GB, of which only ~620 MB were textures and buffers.Once that was fixed by building pipelines at first use, the next wall was the start of gameplay.
ExecuteCommandListswent from 25 ms to 273 ms to 2.3 s per frame with the GPU 2-7 % busy (Metal HUD: 100 % of frame time encoding), and the game's watchdog stopped it.Cause
CreateGraphicsPipelineState/CreateComputePipelineStatecallnewRenderPipelineState/newComputePipelineStateimmediately. Metal compiles each pipeline to GPU code and allocates for it at creation, even though a scene draws only a small fraction of them.mad_convert_stage_optscreates a newMTLLibrary+MTLFunctionfor every stage of every pipeline. The game's 30,370 libraries came from only ~11,400 distinct converter outputs, and Metal keeps GPU-side storage per library.Change
All in
madeira-d3d12/src/pe/madeira_d3d12.c:WMTRenderPipelineInfo(and vertex descriptor) and is built bymad_pso_realizeat its first draw. A compute pipeline is built bymad_cpso_realizeat its first dispatch. Each pipeline has its ownSRWLOCK(zeroed bycalloc), so builds of different pipelines run concurrently. If Metal rejects a pipeline at that point, the failure is logged (first 8) and its draws/dispatches are skipped, like the existing placeholder pipelines, instead ofCreateGraphicsPipelineStatereturningE_FAIL. GS and tessellation pipelines are unchanged (eager).mad_ecl_run,mad_prebuild_listscollects the not-yet-built pipelines that the batch's lists bind (up to 256, skipping lists the replay will reject). When there are at least 2, it builds them on the calling thread plus a persistent pool of 3 worker threads (256 KB stacks, created once). Creating Wine threads per batch cost an 8 MB stack, a TEB and an emulator thread state each time. Batches use the pool one at a time.mad_convert_stage_optshashes the converted metallib bytes plus the entry name (128-bit key). If an identical library exists, it hands out the existingMTLLibrary/MTLFunctionwith a retain of each, sopso_Releaseandmad_tess_freestay balanced. The open-addressing table keeps one reference per distinct library, grows at 70 % load and never fills.pso-lazy = 0restores eager creation, andpso-parallel = 0turns the pre-replay parallel build off. Both are kept because they change when pipeline creation (and its failures) happen, which is useful for A/B runs.ConfigCatalog.generated.swiftgains the two entries exactly asbuild/tools/gen-config-catalog.pyemits them. Only those two lines are added, because the generated file onmainis already stale for unrelated keys.pipelines are built ... (madeira.cfg pso-lazy)oncelazy pipelines built at first draw: Nevery 500pso-parallel: built N pipelines on T threads before replay(first 8 batches, then batches of 32 or more)shared shader libraries: N reuses, M distinctevery 2000Evidence
From the fork, on an iPhone 17 Pro Max (iOS 27) with Ghost of Tsushima:
ExecuteCommandListsper frame (45-48 FPS), withpso-parallel: built 20..69 pipelines on 4 threads. The per-batch thread version caused a storm of thread setups and failed 8 MB reserves at gameplay start. The persistent pool removed it.MTLBinaryArchiveyet).Notes / risks
app/Madeira/arm64ec-windows/madeira_d3d12.dll(andd3d12.dll) is not rebuilt in this PR. It needs a rebuild withbuild/madeira-d3d12/build-pe.sh. I could not run llvm-mingw here. The file was syntax-checked with clang against mingw-w64 headers: no new errors or warnings compared withmain, and a canary edit confirmed the new code is checked.E_FAIL.pso-lazy = 0restores the old contract.p->rps/p->cpsoutside the lock on the fast path. They are written once, under the per-pipeline lock.pso time ...), which are diagnostics only.build/ntdll-unix/virtual_ios.c), which is a separate memory concern.🤖 Generated with Claude Code
Claude-Session: https://claude.ai/code/session_0189oLHghpaYKLk4f786a6bc