Skip to content

perf: cut the processor time of hardware decoding on Windows - #326

Merged
devopvoid merged 2 commits into
mainfrom
perf/d3d12-hardware-decoding
Oct 4, 2026
Merged

devopvoid merged 2 commits into
mainfrom
perf/d3d12-hardware-decoding

Conversation

@devopvoid

Copy link
Copy Markdown
Owner

Stacked on #325 (which is on #323): the base is its branch, so the diff is these two commits only. Retarget to main once those are merged.

Hardware decoding in the media player saved little processor time on Windows, and at 4K cost more than software. Profiling the decoding thread (cycles per picture, QueryThreadCycleTime) showed why, and led to two changes.

1. Reuse the buffers of hardware-decoded frames. For every picture the hardware path dropped the system-memory frame it read the picture back into, and allocated a new I420 frame for the conversion: two fresh allocations of 12 MB per picture at 4K, and a fresh allocation costs more to fault in than the copy into it.

  • The read-back frame now stays between pictures. It gets a new buffer when the picture size, the pixel format or the size of the decoder's surfaces changes. The last matters because the Direct3D 12 download copies the whole padded surface, where Direct3D 11 copies only the smaller of the two frames.
  • The I420 frames come from a buffer pool. The conversion of other formats in software uses it too.
  • Cycles on the decoding thread per picture: read-back 11.0 to 5.3 million at 4K and 3.3 to 1.8 at 1080p; conversion 10.4 to 5.7 and 2.2 to 1.3.

2. Direct3D 12 first on Windows. FFmpeg 8.1 has h264_d3d12va and vp9_d3d12va. The Windows build enables them and DeviceTypes() tries a Direct3D 12 device before Direct3D 11. A machine without Windows 10 2004 or a driver that decodes through Direct3D 12 goes on with Direct3D 11 as before. The component list changed, so the next Windows build rebuilds FFmpeg.

Process CPU per frame in the player on an NVIDIA GeForce RTX 3080, median of four runs of steady playback:

software Direct3D 11 Direct3D 12
1080p H.264 3.6 ms 4.2 ms 2.1 ms
4K H.264 6.6 ms 12.4 ms 5.4 ms

Direct3D 12 uses about 40% less than software at 1080p and 18% less at 4K, and about half of what Direct3D 11 does. The 7 to 8 times of the Apple M2 does not carry over. One machine, synthetic media and large run-to-run noise (the four runs of a setting differ by up to a factor of 2), so these are indications, not a promise; the guide says so and gives these figures.

The pictures are unchanged: all 67 tests of webrtc-java-media pass with webrtc.test.hardwareDecoding set, including the 1080p, 4K, 1366x768 and fallback cases of #325 that compare every frame with software decoding. A log line confirmed the Direct3D 12 device was the one created; it is not part of the change.

Not done:

  • -Xcheck:jni on the media module.
  • Other GPUs and Windows versions. The Windows ARM64 build is untested.
  • --enable-d3d12va makes FFmpeg's configure fail where the Windows SDK has no d3d12video.h. The CI runners most likely have it, but I have not checked, so watch the Windows jobs.

@devopvoid
devopvoid force-pushed the perf/d3d12-hardware-decoding branch from d37a15e to 8378072 Compare October 4, 2026 18:52
@devopvoid
devopvoid added this pull request to stack #327 October 4, 2026 19:10
Base automatically changed from test/hardware-decoding-sizes to main October 4, 2026 19:10
For every picture the hardware path dropped the system-memory frame the
picture was read back into, and allocated a new I420 frame for the
conversion. At 4K that is two fresh allocations of 12 MB per picture, and a
fresh allocation costs more to fault in than the copy into it.

The read-back frame now stays between pictures, with a new buffer only when
the size or the pixel format changes (the transfer copies no more than the
smaller of the two frames), and the I420 frames come from a buffer pool, so a
picture lands in memory that is already mapped. The conversion of other
formats in software uses the pool as well.

CPU cycles on the decoding thread per picture, on an NVIDIA GeForce RTX 3080
through Direct3D 11: the read-back went from 11.0 to 5.3 million at 4K and
from 3.3 to 1.8 million at 1080p, the conversion from 10.4 to 5.7 and from 2.2
to 1.3.
FFmpeg 8.1 has Direct3D 12 hwaccels for H.264 and VP9. The Windows build now
enables them (h264_d3d12va, vp9_d3d12va), and the player tries a Direct3D 12
device before the Direct3D 11 one it used. A machine without Windows 10 2004
or a driver that decodes through Direct3D 12 has no such device, and goes on
with Direct3D 11 as before. The pictures are the same: every test of the
hardware path compares them with software decoding.

Reading a picture back is what costs processor time on this path, and
Direct3D 12 does it for about half of what Direct3D 11 does. Process CPU per
frame in the player on an NVIDIA GeForce RTX 3080, median of four runs of
steady playback, one machine and synthetic media:

                 software   Direct3D 11   Direct3D 12
  1080p H.264      3.6 ms        4.2 ms        2.1 ms
  4K H.264         6.6 ms       12.4 ms        5.4 ms

Direct3D 11 used more processor time than software at 4K, Direct3D 12 less.

The component list changed, so the next build rebuilds FFmpeg on Windows.
@devopvoid
devopvoid force-pushed the perf/d3d12-hardware-decoding branch from 8378072 to ffc55dd Compare October 4, 2026 19:10
@devopvoid
devopvoid merged commit 58253b0 into main Oct 4, 2026
24 checks passed
@devopvoid
devopvoid deleted the perf/d3d12-hardware-decoding branch October 4, 2026 19:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant