Skip to content

feat: decode H.264 and VP9 with NVDEC on NVIDIA GPUs on Linux - #321

Merged
devopvoid merged 2 commits into
mainfrom
feat/linux-nvdec-decoder
Oct 4, 2026
Merged

devopvoid merged 2 commits into
mainfrom
feat/linux-nvdec-decoder

Conversation

@devopvoid

@devopvoid devopvoid commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Important

Not built or run on Linux. It was written on a Mac and compiled and linked on Linux only by CI. The decoder logic has run on an NVIDIA GPU (GeForce RTX 3080), through a temporary Windows build of the same sources (see Testing); what has not been tried is everything specific to Linux: loading the driver libraries, the packaging, and ARM64. It still needs a run on Linux with an NVIDIA GPU: see Needs testing below.

HardwareVideoDecoderFactory now decodes on Linux where there is an NVIDIA GPU: H.264 in all its profiles and VP9 in profile 0, through the NVDEC engine. It is the decoder counterpart of the NVENC encoder (#312), and loads the same way. The Java API is unchanged, and so is what gets negotiated.

Platform HardwareVideoDecoderFactory
Windows H.264, AV1 and VP9 on the GPU, through Media Foundation (#315, #319)
Linux H.264 and VP9 (profile 0) on NVIDIA GPUs, through NVDEC; VA-API (Intel, AMD) is not there yet
macOS H.264 and VP9 through VideoToolbox (#318)

Behavior

  • Loading: libcuda.so.1 and libnvcuvid.so.1 are opened with dlopen when the factory is made, so a machine without the NVIDIA driver is not affected and decodes in software as before. The factory is offered only where the first CUDA device decodes at least one of the codecs (cuvidGetDecoderCaps, 8 bit 4:2:0 into NV12).
  • Negotiation: the formats are the ones the software factory offers (H.264 in all its profiles, VP9 profile 0), so negotiation does not change.
  • Fallback: FallbackVideoDecoder hands the stream to the software decoder when NVDEC returns WEBRTC_VIDEO_CODEC_FALLBACK_SOFTWARE. That happens for another profile or bit depth, a size the GPU does not decode (checked against the caps of the device), a decoder that fails, and VP9 frames with spatial layers: their layers reach a decoder without the superframe index that says where each ends, and NVDEC is not known to take that, where libvpx does. A first frame that is not a key frame asks WebRTC for one.
  • Diagnostics: the decoder shows in the decoderImplementation stat of inbound-rtp as NVDEC (<GPU name>).

Native side

  • NvdecVideoDecoder uses NVDEC's own parser, which finds the pictures in an encoded image and calls back for the sequence (which creates the decoder), each picture to decode, and each to show. With no display delay all of that happens inside Decode, on the caller's thread. A decoded picture is NV12 in GPU memory; it is copied to system memory with cuMemcpy2D, cropped to the picture within the surface (1088 rows for 1080), and converted to I420 with libyuv, as MFVideoDecoder does.
    A frame is matched to its input by a timestamp counted per image, so the decoder does not depend on frames coming out in input order. A resolution change at a key frame makes the decoder again.
  • NvdecLibrary loads the two libraries once, probes the device, and keeps the primary CUDA context, the way NvencLibrary does. NvdecContextScope makes it current for the calls.
  • NvdecVideoDecoderFactory and the Linux platform hook (LinuxHardwareVideoCodecFactories.cpp), which had no decoders.
  • Headers: dependencies/nvdec has dynlink_cuda.h, dynlink_cuviddec.h and dynlink_nvcuvid.h of nv-codec-headers at tag n12.0.16.1, unchanged: the version of the NVENC header, and the same files webrtc-java-media carries for FFmpeg (feat: decode video in hardware in the media player #320). They define the structures NVDEC's parser and decoder take, which would be error-prone to declare by hand. They are compile-time only. Their licenses are gathered in LICENSE and installed into the Linux platform jars under META-INF/licenses/nvdec.
  • CMake: the NVDEC sources and include path are for Linux only; Windows still decodes through Media Foundation.
  • Left out on purpose: AV1, until it can be tried on an RTX 30 or newer; VA-API for Intel and AMD GPUs, which needs the slice and frame headers parsed by hand.

Testing

  • HardwareVideoDecoderIntegrationTest runs on Linux too: the H.264 test and the VP9 tests (hardwareDecodesVp9, hardwareFollowsResolutionChange, vp9NeedsNoKeyFrames, vp9TemporalLayers, vp9SpatialLayers) check there what they check on Windows, and the implementation name NVDEC counts as hardware. A decoder is required with -Dwebrtc.test.hardwareDecoder=true (H.264) and -Dwebrtc.test.hardwareVp9Decoder=true (VP9); without them the tests are skipped where there is no decoder.
  • On an Apple M2 (macOS 14.5) the build and the decoder tests run as before (8 passed, 2 skipped). The new sources were compiled with clang++ -std=c++20 -Wall -Wextra -fsyntax-only against the project's WebRTC headers with Linux defines, with a negative control to show that the check reports errors. The only warning is the unused env parameter that the NVENC factory has too.
  • CI builds and links the sources on all three Linux targets (x86-64, ARM64, ARM) and the tests pass there. The runners have no GPU, so the hardware tests are skipped.
  • Run on an NVIDIA GPU, through Windows. NVDEC's cuvid* API is the same on Windows, so the decoder was built there, locally and not committed, with nvcuda.dll and nvcuvid.dll in place of the Linux library names, the NVDEC factory ahead of Media Foundation, and isHardware accepting NVDEC only, so that a fallback to Media Foundation would fail a test. On a GeForce RTX 3080:
    • all 13 tests of HardwareVideoDecoderIntegrationTest pass (3 skipped: the macOS one and the two AV1 ones), including the padded-frame crop tests for H.264 and VP9 (640x360) and vp9NeedsNoKeyFrames;
    • the decoder name is NVDEC (NVIDIA GeForce RTX 3080) for VP9 at 128x128, 320x240 and 641x361, and for H.264 at 160x120, 320x240 and 1920x1080;
    • a resolution change from 640x480 to 320x240 stays on NVDEC, for VP9 and H.264;
    • 15 calls in a row all got NVDEC (no session or context leaks), and four mixed H.264 and VP9 streams at once all did;
    • no decode failure was logged.
  • A test fix from that run: hardwareFollowsResolutionChange halved 320x240 to 160x120 and asserted that the decoder was still hardware. NVDEC does not decode VP9 below 128 pixels on the shorter side (the log reads does not decode 160x120), so it refuses that size, as it should, and the stream goes on in the software decoder. The test now sends 640x480 and halves it to 320x240.

Not verified yet:

  • Everything specific to Linux: dlopen of libcuda.so.1 and libnvcuvid.so.1, and the packaging of the licenses into the platform jars.
  • ARM64 Linux with an NVIDIA GPU, which CI builds but cannot run.
  • Streams that reorder frames: the frames are matched by timestamp, but nothing here sends B-frames.
  • cuvidCtxLock is not used. Several decoders on the one context ran fine in the test above, but that was one process on one GPU and driver.

Known nits, not changed: NvdecLibrary::Supports counts macroblocks with width / 16 * height / 16, which rounds down and is only wrong at the limit of the GPU; and the sequence callback makes the decoder again even when the stream's parameters have not changed.

Needs testing

On Linux with an NVIDIA GPU and its driver (libcuda.so.1 and libnvcuvid.so.1 present):

mvn -pl webrtc test -Dtest=HardwareVideoDecoderIntegrationTest -Dwebrtc.test.hardwareDecoder=true -Dwebrtc.test.hardwareVp9Decoder=true

This fails unless H.264 and VP9 are decoded in hardware. The decoderImplementation stat of a call should read NVDEC (<GPU name>). The log line NVDEC available on <GPU>, H.264: .., VP9: .. tells what the device offers. What the Windows run could not show, and is worth a look on Linux: that the two libraries are found, and a call with several video streams at once on that driver.

Docs

The video codecs guide (docs/guide/advanced/video-codecs.md) has NVDEC in the Linux row and says what it needs. The Javadoc of HardwareVideoDecoderFactory has the same.

HardwareVideoDecoderFactory now decodes on Linux where there is an NVIDIA
GPU: H.264 in all its profiles and VP9 in profile 0, through the NVDEC
engine. It is the decoder counterpart of the NVENC encoder, and loads the
same way: libcuda.so.1 and libnvcuvid.so.1 are opened at run time, so a
machine without the NVIDIA driver is not affected, and falls back to the
software decoders as before. What gets negotiated does not change.

NvdecVideoDecoder uses NVDEC's own parser, which finds the pictures in an
encoded image and calls back for the sequence (which creates the decoder),
each picture to decode, and each to show. With no display delay all of that
happens inside Decode. A decoded picture is NV12 in GPU memory; it is copied
to system memory, cropped to the picture within the surface, and converted
to I420 with libyuv, as the Media Foundation decoder does on Windows.

Whatever NVDEC cannot take returns WEBRTC_VIDEO_CODEC_FALLBACK_SOFTWARE and
FallbackVideoDecoder hands the stream to the software decoder: another
profile or bit depth, a size the GPU does not decode, a failure of the
decoder, and VP9 frames with spatial layers, whose layers reach a decoder
without the index that says where each ends. A first frame that is not a key
frame asks WebRTC for one.

The headers are the three of nv-codec-headers (tag n12.0.16.1, the version
of the NVENC header) that NVDEC needs, unchanged, in dependencies/nvdec with
their licenses; they are compile-time only. They are the same files
webrtc-java-media carries for FFmpeg.

The decoder tests of the webrtc module run on Linux too: H.264 and VP9 are
checked there as on Windows, and the implementation name NVDEC counts as
hardware.

This has not been built or run on Linux, and never on an NVIDIA GPU. The
sources compile with clang against WebRTC's headers on a Mac (syntax only,
with Linux defines); how NVDEC behaves on real streams could not be seen.
@devopvoid
devopvoid force-pushed the feat/linux-nvdec-decoder branch from 7ef5a99 to d84ba77 Compare October 4, 2026 09:42
hardwareFollowsResolutionChange halved 320x240 to 160x120 and asserted that
the decoder was still a hardware one. NVDEC does not decode VP9 below 128
pixels on the shorter side, so it refuses that size, as it should, and the
stream goes on in libvpx; the test failed on an NVIDIA GPU for that reason.

The call now sends 640x480 and halves it to 320x240, which every hardware
decoder takes.

Run on an NVIDIA GeForce RTX 3080 with NVDEC built on Windows, which is the
same decoder code over the same cuvid API: the test passes with NVDEC the
only decoder accepted as hardware.
@devopvoid
devopvoid merged commit 6bd0492 into main Oct 4, 2026
24 checks passed
@devopvoid
devopvoid deleted the feat/linux-nvdec-decoder branch October 4, 2026 18:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant