Skip to content

Native D3D12: build pipelines at first use, in parallel, and share id… - #139

Merged
willfaust merged 1 commit into
willfaust:mainfrom
bahacan16:pr/d3d12-pipeline-build
Oct 4, 2026
Merged

willfaust merged 1 commit into
willfaust:mainfrom
bahacan16:pr/d3d12-pipeline-build

Conversation

@bahacan16

Copy link
Copy Markdown
Contributor

…entical libraries

Branch: pr/d3d12-pipeline-build (one commit). Companion PR: pr/d3d12-persistent-shader-cache.

Problem

Ghost of Tsushima (Nixxes port) creates every pipeline of a level up front: about 14,000 at New Game, with 29,297 shader libraries. On an iPhone 17 Pro Max the loading screen spun until iOS killed the app. The last footprint line read 7.8 GB, with Metal's currentAllocatedSize at 5.1 GB, of which only ~620 MB were textures and buffers.

Once that was fixed by building pipelines at first use, the next wall was the start of gameplay. ExecuteCommandLists went from 25 ms to 273 ms to 2.3 s per frame with the GPU 2-7 % busy (Metal HUD: 100 % of frame time encoding), and the game's watchdog stopped it.

Cause

  • CreateGraphicsPipelineState / CreateComputePipelineState call newRenderPipelineState / newComputePipelineState immediately. Metal compiles each pipeline to GPU code and allocates for it at creation, even though a scene draws only a small fraction of them.
  • mad_convert_stage_opts creates a new MTLLibrary + MTLFunction for every stage of every pipeline. The game's 30,370 libraries came from only ~11,400 distinct converter outputs, and Metal keeps GPU-side storage per library.
  • Built lazily, pipelines compile on the submitting thread, one after another, at the first draw. The first gameplay frames need hundreds.

Change

All in madeira-d3d12/src/pe/madeira_d3d12.c:

  • Lazy pipelines. A plain render pipeline (no GS, no hull/domain) keeps its WMTRenderPipelineInfo (and vertex descriptor) and is built by mad_pso_realize at its first draw. A compute pipeline is built by mad_cpso_realize at its first dispatch. Each pipeline has its own SRWLOCK (zeroed by calloc), so builds of different pipelines run concurrently. If Metal rejects a pipeline at that point, the failure is logged (first 8) and its draws/dispatches are skipped, like the existing placeholder pipelines, instead of CreateGraphicsPipelineState returning E_FAIL. GS and tessellation pipelines are unchanged (eager).
  • Parallel first use. At the start of mad_ecl_run, mad_prebuild_lists collects the not-yet-built pipelines that the batch's lists bind (up to 256, skipping lists the replay will reject). When there are at least 2, it builds them on the calling thread plus a persistent pool of 3 worker threads (256 KB stacks, created once). Creating Wine threads per batch cost an 8 MB stack, a TEB and an emulator thread state each time. Batches use the pool one at a time.
  • Shared Metal libraries. mad_convert_stage_opts hashes the converted metallib bytes plus the entry name (128-bit key). If an identical library exists, it hands out the existing MTLLibrary/MTLFunction with a retain of each, so pso_Release and mad_tess_free stay balanced. The open-addressing table keeps one reference per distinct library, grows at 70 % load and never fills.
  • Switches (madeira.cfg, both default 1): pso-lazy = 0 restores eager creation, and pso-parallel = 0 turns the pre-replay parallel build off. Both are kept because they change when pipeline creation (and its failures) happen, which is useful for A/B runs. ConfigCatalog.generated.swift gains the two entries exactly as build/tools/gen-config-catalog.py emits them. Only those two lines are added, because the generated file on main is already stale for unrelated keys.
  • Log lines, in the file's existing style:
    • pipelines are built ... (madeira.cfg pso-lazy) once
    • lazy pipelines built at first draw: N every 500
    • pso-parallel: built N pipelines on T threads before replay (first 8 batches, then batches of 32 or more)
    • shared shader libraries: N reuses, M distinct every 2000

Evidence

From the fork, on an iPhone 17 Pro Max (iOS 27) with Ghost of Tsushima:

  • With lazy pipelines plus shared libraries (and the fork's persistent shader cache), New Game no longer runs out of memory. About 500 pipelines were built at first draw in the first scene, and the game reaches gameplay (~20-40 FPS at that point).
  • Parallel prebuild: the menu dropped to ~2 ms of ExecuteCommandLists per frame (45-48 FPS), with pso-parallel: built 20..69 pipelines on 4 threads. The per-batch thread version caused a storm of thread setups and failed 8 MB reserves at gameplay start. The persistent pool removed it.
  • Remaining cost: a pipeline used for the first time in a new area still causes a 50-550 ms hitch while Metal compiles it (no MTLBinaryArchive yet).

Notes / risks

  • The prebuilt app/Madeira/arm64ec-windows/madeira_d3d12.dll (and d3d12.dll) is not rebuilt in this PR. It needs a rebuild with build/madeira-d3d12/build-pe.sh. I could not run llvm-mingw here. The file was syntax-checked with clang against mingw-w64 headers: no new errors or warnings compared with main, and a canary edit confirmed the new code is checked.
  • The shared-library table keeps every distinct library alive for the life of the process, even after all pipelines using it are released. That was the right trade for this game (11,400 shared vs 30,370 separate), but a title that streams many unique pipelines and frees them would hold them longer than before.
  • A pipeline that Metal rejects is now discovered at first draw rather than at creation, so the app gets a pipeline object whose draws are skipped instead of E_FAIL. pso-lazy = 0 restores the old contract.
  • The lazy build reads p->rps/p->cps outside the lock on the fast path. They are written once, under the per-pipeline lock.
  • Not ported:
    • The fork's per-phase timing counters (pso time ...), which are diagnostics only.
    • The ntdll "wider high-band fallback" that shared a fork commit with the thread pool (build/ntdll-unix/virtual_ios.c), which is a separate memory concern.
    • The placeholder for hull+domain pipelines that cannot be built, which belongs with tessellation.
  • No submodule changes.

🤖 Generated with Claude Code

Claude-Session: https://claude.ai/code/session_0189oLHghpaYKLk4f786a6bc

…entical libraries

Branch: `pr/d3d12-pipeline-build` (one commit). Companion PR: `pr/d3d12-persistent-shader-cache`.

## Problem

Ghost of Tsushima (Nixxes port) creates every pipeline of a level up front: about 14,000 at New Game, with 29,297 shader libraries. On an iPhone 17 Pro Max the loading screen spun until iOS killed the app. The last footprint line read 7.8 GB, with Metal's `currentAllocatedSize` at 5.1 GB, of which only ~620 MB were textures and buffers.

Once that was fixed by building pipelines at first use, the next wall was the start of gameplay. `ExecuteCommandLists` went from 25 ms to 273 ms to 2.3 s per frame with the GPU 2-7 % busy (Metal HUD: 100 % of frame time encoding), and the game's watchdog stopped it.

## Cause

- `CreateGraphicsPipelineState` / `CreateComputePipelineState` call `newRenderPipelineState` / `newComputePipelineState` immediately. Metal compiles each pipeline to GPU code and allocates for it at creation, even though a scene draws only a small fraction of them.
- `mad_convert_stage_opts` creates a new `MTLLibrary` + `MTLFunction` for every stage of every pipeline. The game's 30,370 libraries came from only ~11,400 distinct converter outputs, and Metal keeps GPU-side storage per library.
- Built lazily, pipelines compile on the submitting thread, one after another, at the first draw. The first gameplay frames need hundreds.

## Change

All in `madeira-d3d12/src/pe/madeira_d3d12.c`:

- **Lazy pipelines.** A plain render pipeline (no GS, no hull/domain) keeps its `WMTRenderPipelineInfo` (and vertex descriptor) and is built by `mad_pso_realize` at its first draw. A compute pipeline is built by `mad_cpso_realize` at its first dispatch. Each pipeline has its own `SRWLOCK` (zeroed by `calloc`), so builds of different pipelines run concurrently. If Metal rejects a pipeline at that point, the failure is logged (first 8) and its draws/dispatches are skipped, like the existing placeholder pipelines, instead of `CreateGraphicsPipelineState` returning `E_FAIL`. GS and tessellation pipelines are unchanged (eager).
- **Parallel first use.** At the start of `mad_ecl_run`, `mad_prebuild_lists` collects the not-yet-built pipelines that the batch's lists bind (up to 256, skipping lists the replay will reject). When there are at least 2, it builds them on the calling thread plus a persistent pool of 3 worker threads (256 KB stacks, created once). Creating Wine threads per batch cost an 8 MB stack, a TEB and an emulator thread state each time. Batches use the pool one at a time.
- **Shared Metal libraries.** `mad_convert_stage_opts` hashes the converted metallib bytes plus the entry name (128-bit key). If an identical library exists, it hands out the existing `MTLLibrary`/`MTLFunction` with a retain of each, so `pso_Release` and `mad_tess_free` stay balanced. The open-addressing table keeps one reference per distinct library, grows at 70 % load and never fills.
- Switches (madeira.cfg, both default 1): `pso-lazy = 0` restores eager creation, and `pso-parallel = 0` turns the pre-replay parallel build off. Both are kept because they change when pipeline creation (and its failures) happen, which is useful for A/B runs. `ConfigCatalog.generated.swift` gains the two entries exactly as `build/tools/gen-config-catalog.py` emits them. Only those two lines are added, because the generated file on `main` is already stale for unrelated keys.
- Log lines, in the file's existing style:
  - `pipelines are built ... (madeira.cfg pso-lazy)` once
  - `lazy pipelines built at first draw: N` every 500
  - `pso-parallel: built N pipelines on T threads before replay` (first 8 batches, then batches of 32 or more)
  - `shared shader libraries: N reuses, M distinct` every 2000

## Evidence

From the fork, on an iPhone 17 Pro Max (iOS 27) with Ghost of Tsushima:

- With lazy pipelines plus shared libraries (and the fork's persistent shader cache), New Game no longer runs out of memory. About 500 pipelines were built at first draw in the first scene, and the game reaches gameplay (~20-40 FPS at that point).
- Parallel prebuild: the menu dropped to ~2 ms of `ExecuteCommandLists` per frame (45-48 FPS), with `pso-parallel: built 20..69 pipelines on 4 threads`. The per-batch thread version caused a storm of thread setups and failed 8 MB reserves at gameplay start. The persistent pool removed it.
- Remaining cost: a pipeline used for the first time in a new area still causes a 50-550 ms hitch while Metal compiles it (no `MTLBinaryArchive` yet).

## Notes / risks

- **The prebuilt `app/Madeira/arm64ec-windows/madeira_d3d12.dll` (and `d3d12.dll`) is not rebuilt in this PR.** It needs a rebuild with `build/madeira-d3d12/build-pe.sh`. I could not run llvm-mingw here. The file was syntax-checked with clang against mingw-w64 headers: no new errors or warnings compared with `main`, and a canary edit confirmed the new code is checked.
- The shared-library table keeps every distinct library alive for the life of the process, even after all pipelines using it are released. That was the right trade for this game (11,400 shared vs 30,370 separate), but a title that streams many unique pipelines and frees them would hold them longer than before.
- A pipeline that Metal rejects is now discovered at first draw rather than at creation, so the app gets a pipeline object whose draws are skipped instead of `E_FAIL`. `pso-lazy = 0` restores the old contract.
- The lazy build reads `p->rps`/`p->cps` outside the lock on the fast path. They are written once, under the per-pipeline lock.
- **Not ported:**
  - The fork's per-phase timing counters (`pso time ...`), which are diagnostics only.
  - The ntdll "wider high-band fallback" that shared a fork commit with the thread pool (`build/ntdll-unix/virtual_ios.c`), which is a separate memory concern.
  - The placeholder for hull+domain pipelines that cannot be built, which belongs with tessellation.
- No submodule changes.

---
🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0189oLHghpaYKLk4f786a6bc
Signed-off-by: bahacan16 <190844990+bahacan16@users.noreply.github.com>
willfaust pushed a commit that referenced this pull request Oct 4, 2026
#139)

Squashed from #139.

Signed-off-by: bahacan16 <190844990+bahacan16@users.noreply.github.com>
willfaust pushed a commit that referenced this pull request Oct 4, 2026
…on (#159)

Squashed from #159, on top of #139's lazy pipelines.

Signed-off-by: bahacan16 <190844990+bahacan16@users.noreply.github.com>
willfaust added a commit that referenced this pull request Oct 4, 2026
d3d12.dll and madeira_d3d12.dll rebuilt from #138, #139, #152-#159, #165,
#166 and the fixes on top; d3d12core.dll (#166) is new and is used only
with madeira.cfg d3d12-core-dll = 1.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@willfaust willfaust closed this Oct 4, 2026
@willfaust willfaust reopened this Oct 4, 2026
@willfaust
willfaust merged commit a70caf9 into willfaust:main Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants